• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Inworld AI Announces Realtime TTS-2 General Availability

    Robert Scoble retweeted Inworld AI's announcement of the model's release.

    RS
    CH
    AA
    12 Sources, 28d ago, first seen 28d ago

    TLDR

    Robert Scoble retweeted a post from the @inworld_ai account on X. The retweeted post claims that Realtime TTS-2 reached general availability today and is now the #1 model on Artificial Analysis. It also claims the model is the fastest text-to-speech system in its class. The cluster contains only this single retweet, with no additional details or corroboration from other sources.

    Combined views

    935.8K

    12 Sources, first seen 28d ago

    Combined views

    935.8K

    12 Sources, first seen 28d ago

    1.8K likes
    1.8K likes
    196 comments
    1.4K saves
    323 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    196 comments
    1.4K saves
    323 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    12 Sources

    @inworldRealtime TTS-2 is GA today, now the #1 model on Artificial Analysis and fastest TTS in the world of its class. The research preview got far more usage than expected. That experience went back into research and today's GA model represents the biggest leap in Inworld's history.
    @ScobleizerRT @inworld_ai: Realtime TTS-2 is GA today, now the #1 model on Artificial Analysis and fastest TTS in the world of its class. The researc…
    @kimmonismusInworld’s Realtime TTS-2 is now generally available and currently ranks #1 in Artificial Analysis’ Controlled Voice Arena. Developers can write delivery instructions into the text, from a short emotion tag to a full sentence describing the tone.
    @svpino~100ms latency is really fast for a Text-To-Speech model. For reference, a friend's company uses a voice model in production with average latency just over 300ms. They process around 7,500 messages every day using that model.
    @ArtificialAnlysInworld's newly released Realtime TTS-2 is the new #1 on the Artificial Analysis Controlled Voice Arena, just ahead of Cartesia Sonic 3.6, and ranks #2 on our Provider Voice Arena behind Sonic 3.6 Realtime TTS-2 is a new Text to Speech model from @inworld_ai that supports over 100 languages, including English, Hindi, Spanish, French, German, Chinese, and Japanese. Language can be set explicitly or detected from the text, including multiple languages in a single request (e.g., "I'll grab a coffee. ¿Quieres uno? お疲れさま。"). It also supports delivery instructions written in plain text alongside the input (e.g., [speak tired but warm, like she just got home from a long day]). Key takeaways: ➤ Controlled Voice: Realtime TTS-2 takes #1 on the Controlled Voice Arena with an Elo of 1,123 (+16/-16) across 1,292 appearances, 4 points ahead of Cartesia's Sonic 3.6 at 1,119. Sonic 3.5 follows at 1,096, then Inworld’s own Realtime TTS-2 Flash - Research Preview at 1,075, and ElevenLabs Eleven v3 at 1,062 ➤ Provider Voice: Realtime TTS-2 takes #2 on the Provider Voice Arena with an Elo score of 1,252 (+18/-18) based on 1,094 arena appearances, placing it ahead of Alibaba Qwen-Audio-3.0-TTS-Plus at 1,241 and Speechify Simba 3.2 at 1,240, but behind Sonic 3.6 at 1,282 ➤ Throughput: The model processes 106 characters per second of generation time, compared to 124 for Sonic 3.6, 97 for Simba 3.2, and 40 for Eleven v3 ➤ Pricing: Realtime TTS-2 is priced at $20.83 per 1M characters, lower priced than Sonic 3.6 ($49), Eleven v3 ($100), and Qwen-Audio-3.0-TTS-Plus ($27.59), but more expensive than Simba 3.2 ($10) See more details and listen to samples in the thread below ⬇️

    12 Sources

    @inworldRealtime TTS-2 is GA today, now the #1 model on Artificial Analysis and fastest TTS in the world of its class. The research preview got far more usage than expected. That experience went back into research and today's GA model represents the biggest leap in Inworld's history.
    @ScobleizerRT @inworld_ai: Realtime TTS-2 is GA today, now the #1 model on Artificial Analysis and fastest TTS in the world of its class. The researc…
    @kimmonismusInworld’s Realtime TTS-2 is now generally available and currently ranks #1 in Artificial Analysis’ Controlled Voice Arena. Developers can write delivery instructions into the text, from a short emotion tag to a full sentence describing the tone.
    @svpino~100ms latency is really fast for a Text-To-Speech model. For reference, a friend's company uses a voice model in production with average latency just over 300ms. They process around 7,500 messages every day using that model.
    @ArtificialAnlysInworld's newly released Realtime TTS-2 is the new #1 on the Artificial Analysis Controlled Voice Arena, just ahead of Cartesia Sonic 3.6, and ranks #2 on our Provider Voice Arena behind Sonic 3.6 Realtime TTS-2 is a new Text to Speech model from @inworld_ai that supports over 100 languages, including English, Hindi, Spanish, French, German, Chinese, and Japanese. Language can be set explicitly or detected from the text, including multiple languages in a single request (e.g., "I'll grab a coffee. ¿Quieres uno? お疲れさま。"). It also supports delivery instructions written in plain text alongside the input (e.g., [speak tired but warm, like she just got home from a long day]). Key takeaways: ➤ Controlled Voice: Realtime TTS-2 takes #1 on the Controlled Voice Arena with an Elo of 1,123 (+16/-16) across 1,292 appearances, 4 points ahead of Cartesia's Sonic 3.6 at 1,119. Sonic 3.5 follows at 1,096, then Inworld’s own Realtime TTS-2 Flash - Research Preview at 1,075, and ElevenLabs Eleven v3 at 1,062 ➤ Provider Voice: Realtime TTS-2 takes #2 on the Provider Voice Arena with an Elo score of 1,252 (+18/-18) based on 1,094 arena appearances, placing it ahead of Alibaba Qwen-Audio-3.0-TTS-Plus at 1,241 and Speechify Simba 3.2 at 1,240, but behind Sonic 3.6 at 1,282 ➤ Throughput: The model processes 106 characters per second of generation time, compared to 124 for Sonic 3.6, 97 for Simba 3.2, and 40 for Eleven v3 ➤ Pricing: Realtime TTS-2 is priced at $20.83 per 1M characters, lower priced than Sonic 3.6 ($49), Eleven v3 ($100), and Qwen-Audio-3.0-TTS-Plus ($27.59), but more expensive than Simba 3.2 ($10) See more details and listen to samples in the thread below ⬇️