Inworld's Realtime TTS-2 goes live: 100 ms voice and one identity across 200+ languages

Inworld's Realtime TTS-2 is now generally available, with about 100 ms to first audio, one voice identity across 200+ languages and a 25 ms Flash model at half the price.

Paylaş
Inworld's Realtime TTS-2 goes live: 100 ms voice and one identity across 200+ languages

In its official announcement, Inworld AI made Realtime TTS-2 generally available, calling it the biggest leap in the company's history, and released a faster companion model, TTS-2 Flash, the same day. Inworld reports about 100 ms to the first byte of audio for the main model and 25 ms for Flash, and says a single voice identity can now speak in more than 200 languages. TTS means text to speech: software that turns written text into spoken audio. The company also claims the top spot on the third party ranking site Artificial Analysis and the fastest model in its class.

Directing a voice in plain language

The headline feature is voice direction: performance notes typed in square brackets ahead of the line. Inworld's example runs one sentence, "I'm happy for you. Really.", first after "[speak through gritted teeth, barely holding it in]" and then after "[speak warmly, like you mean every word]". Same words, different acting. The release also adds self-serve professional voice cloning, voice design in plain language, cross-lingual output that keeps the voice identity, and multi-turn context. Two days before the launch, Inworld open sourced its TTS evaluation kit, noting that two teams running the same model on the same data can report different word error rates depending on the speech recognition model used to score them.

Speed, price and where to get it

On a subscription Realtime TTS-2 costs 12.50 dollars per million characters, roughly 1.5 cents a minute, and Flash is half that. Inworld aims Flash at uses where every millisecond counts and says it outruns models with twenty times the cost and delay. Hundreds of millions of people use apps built on the model every day, the company says, from language learning and tutoring to roleplay, fitness, news, games and customer support. LiveKit CTO David Zhao, quoted in the announcement, called it a step forward in emotionally expressive speech synthesis. The models run at inworld.ai and through partners including Cloudflare, DeepInfra, LiveKit, Pipecat, Telnyx and Tencent RTC.

Devamını oku

Anthropic 开源 Claude Commerce Agents:购物与商家两个智能体,零售、旅行、电信、娱乐四套参考实现

Anthropic 开源 Claude Commerce Agents:购物与商家两个智能体,零售、旅行、电信、娱乐四套参考实现

Anthropic 把搭建电商智能体的参考蓝图开源:购物智能体嵌进商家应用替顾客比价填购物车,商家智能体供员工管库存与定价,零售、旅行、电信、娱乐四个场景可直接跑。架构是一个标准循环里只用一个模型。仓库里的公司和品牌全部虚构,示例不下单不扣卡,金融操作一律等人批准。官方称提示缓存命中率 90% 到 99%、购物车最多大出 35%、成交可能性高出 60%,这些数字均属厂商说法。

GlobalFeed Editor tarafından