Inworld's Realtime TTS-2 goes live: 100 ms voice and one identity across 200+ languages
Inworld's Realtime TTS-2 is now generally available, with about 100 ms to first audio, one voice identity across 200+ languages and a 25 ms Flash model at half the price.
In its official announcement, Inworld AI made Realtime TTS-2 generally available, calling it the biggest leap in the company's history, and released a faster companion model, TTS-2 Flash, the same day. Inworld reports about 100 ms to the first byte of audio for the main model and 25 ms for Flash, and says a single voice identity can now speak in more than 200 languages. TTS means text to speech: software that turns written text into spoken audio. The company also claims the top spot on the third party ranking site Artificial Analysis and the fastest model in its class.
Directing a voice in plain language
The headline feature is voice direction: performance notes typed in square brackets ahead of the line. Inworld's example runs one sentence, "I'm happy for you. Really.", first after "[speak through gritted teeth, barely holding it in]" and then after "[speak warmly, like you mean every word]". Same words, different acting. The release also adds self-serve professional voice cloning, voice design in plain language, cross-lingual output that keeps the voice identity, and multi-turn context. Two days before the launch, Inworld open sourced its TTS evaluation kit, noting that two teams running the same model on the same data can report different word error rates depending on the speech recognition model used to score them.
Speed, price and where to get it
On a subscription Realtime TTS-2 costs 12.50 dollars per million characters, roughly 1.5 cents a minute, and Flash is half that. Inworld aims Flash at uses where every millisecond counts and says it outruns models with twenty times the cost and delay. Hundreds of millions of people use apps built on the model every day, the company says, from language learning and tutoring to roleplay, fitness, news, games and customer support. LiveKit CTO David Zhao, quoted in the announcement, called it a step forward in emotionally expressive speech synthesis. The models run at inworld.ai and through partners including Cloudflare, DeepInfra, LiveKit, Pipecat, Telnyx and Tencent RTC.