Inworld's Realtime TTS-2 goes live: 100 ms voice and one identity across 200+ languages

Inworld's Realtime TTS-2 is now generally available, with about 100 ms to first audio, one voice identity across 200+ languages and a 25 ms Flash model at half the price.

Paylaş
Inworld's Realtime TTS-2 goes live: 100 ms voice and one identity across 200+ languages

In its official announcement, Inworld AI made Realtime TTS-2 generally available, calling it the biggest leap in the company's history, and released a faster companion model, TTS-2 Flash, the same day. Inworld reports about 100 ms to the first byte of audio for the main model and 25 ms for Flash, and says a single voice identity can now speak in more than 200 languages. TTS means text to speech: software that turns written text into spoken audio. The company also claims the top spot on the third party ranking site Artificial Analysis and the fastest model in its class.

Directing a voice in plain language

The headline feature is voice direction: performance notes typed in square brackets ahead of the line. Inworld's example runs one sentence, "I'm happy for you. Really.", first after "[speak through gritted teeth, barely holding it in]" and then after "[speak warmly, like you mean every word]". Same words, different acting. The release also adds self-serve professional voice cloning, voice design in plain language, cross-lingual output that keeps the voice identity, and multi-turn context. Two days before the launch, Inworld open sourced its TTS evaluation kit, noting that two teams running the same model on the same data can report different word error rates depending on the speech recognition model used to score them.

Speed, price and where to get it

On a subscription Realtime TTS-2 costs 12.50 dollars per million characters, roughly 1.5 cents a minute, and Flash is half that. Inworld aims Flash at uses where every millisecond counts and says it outruns models with twenty times the cost and delay. Hundreds of millions of people use apps built on the model every day, the company says, from language learning and tutoring to roleplay, fitness, news, games and customer support. LiveKit CTO David Zhao, quoted in the announcement, called it a step forward in emotionally expressive speech synthesis. The models run at inworld.ai and through partners including Cloudflare, DeepInfra, LiveKit, Pipecat, Telnyx and Tencent RTC.

Devamını oku

Inworld 称 Realtime TTS-2 正式商用:首字节延迟约 100 毫秒,同一音色支持 200 多种语言,同日发布的 Flash 版本降至 25 毫秒

Inworld 称 Realtime TTS-2 正式商用:首字节延迟约 100 毫秒,同一音色支持 200 多种语言,同日发布的 Flash 版本降至 25 毫秒

Inworld 于 9 月 2 日宣布文本转语音模型 Realtime TTS-2 正式商用(GA),并同步推出更快的 TTS-2 Flash。公司称前者首字节延迟约 100 毫秒、同一音色可跨 200 多种语言使用,订阅价为每百万字符 12.5 美元;后者延迟低至 25 毫秒,价格为其一半。模型已在 inworld.ai 上线,也可通过 LiveKit、Cloudflare 等平台调用。

GlobalFeed Editor tarafından
Inworld تطلق Realtime TTS-2 رسميا: 100 مللي ثانية وهوية صوت واحدة في أكثر من 200 لغة

Inworld تطلق Realtime TTS-2 رسميا: 100 مللي ثانية وهوية صوت واحدة في أكثر من 200 لغة

أعلنت Inworld الإتاحة العامة لنموذج Realtime TTS-2 لتحويل النص إلى كلام، بزمن استجابة نحو 100 مللي ثانية ودعم هوية صوتية واحدة عبر أكثر من 200 لغة، إلى جانب إصدار TTS-2 Flash بزمن 25 مللي ثانية ونصف السعر.

GlobalFeed Editor tarafından
World Labs发布Atlas:号称全球首个多模态世界模型,像素级相机控制生成1440p视频、稀疏图片重建3D,可为机器人生成模拟数据,未来数周开放早期访问

World Labs发布Atlas:号称全球首个多模态世界模型,像素级相机控制生成1440p视频、稀疏图片重建3D,可为机器人生成模拟数据,未来数周开放早期访问

李飞飞创办的World Labs于9月1日发布多模态世界模型Atlas,官方称其为全球首个此类模型:可从一张或多张图片以像素级相机控制生成最长1分钟的1440p视频,从少量照片重建真实空间的3D结构,并为机器人生成RGB与深度模拟数据。公司自测在相机条件生成与3D重建基准上超过专用模型。目前仅向选定合作伙伴开放早期访问,未来几周扩大范围,未公布价格。

GlobalFeed Editor tarafından