Meta ships live transcription in a single model: Muse Voice Transcribe enters preview

Muse Voice Transcribe folds real-time speech recognition and live speaker attribution into one streaming model; it is in public preview on the Meta Model API, aimed at developers who do not want to build a pipeline.

Paylaş
Meta ships live transcription in a single model: Muse Voice Transcribe enters preview

The classic ordeal of a developer adding live speech to a product is a pipeline: one model transcribes, a second separates speakers, a third adds punctuation, and their latencies stack. Meta's Muse Voice Transcribe announcement aims straight at that ordeal: real-time transcription and live speaker attribution combined in a single streaming model, now in public preview on the Meta Model API.

Meta also filled in the model's identity card: Muse Voice Transcribe is the first real-time audio perception model from the Superintelligence Labs team. The capabilities sit in one stream: streaming speech recognition, conversational diarization for 20-plus speakers and endpoint detection; it is multilingual, follows mid-sentence language switches seamlessly, and accuracy can be raised with language, keyword and context biasing. Independent measurement backs the claim: on Artificial Analysis's streaming AA-WER index the model ranks first at a 3.1 percent error rate, ahead of Cartesia at 3.4 and ElevenLabs at 3.6.

Meta's claim is competitive transcription accuracy without stitching helper models together and without paying premium prices. The translation: products like meeting assistants, live captions and call-center analytics get both "what was said" and "who said it" from one API call, dropping the cost and latency of a separate diarization layer.

The positioning matters too: as Meta opens its Model API to enterprise developers with the Muse family, it lands on the same shelf as OpenAI's Whisper line and Google's speech services. Voice is the natural interface of the agent era; read alongside the on-device wave we covered this week, the picture sharpens: speech recognition is on its way from separate product to default input of every product.