Tavus's Sparrow-2 dismantles the voice AI pipeline: it listens to the whole room and the whole turn
The conventional system cleans the audio, isolates one speaker and throws away the rest. Sparrow-2 does the opposite, and the company says it measured a 2.1 percent conversation failure rate.
Tavus has taken Sparrow-2, its real-time conversational understanding model, live. The model powers what the company calls PALs, its video-based conversational AI characters.
Understanding the problem means knowing the existing method. Voice AI has so far made the task tractable like this: voice activity detection finds who is speaking, noise cancellation erases the room, a single speaker is isolated, the audio is converted to text, and a timer decides the sentence has ended. Information is discarded at every step. An mhm counts as noise, a second voice becomes interference, and a loud room disappears entirely.
Sparrow-2 does the opposite: it continuously models the whole turn and the whole room from streaming audio. By the company's account this unlocks two things. First, voice AI no longer needs a quiet room and a clean microphone, which opens up kiosks, retail floors, cafés and cars; the matter known in the literature as the cocktail party problem. Second, turn-taking becomes more natural: it distinguishes a pause that means wait, an mhm that means keep going, and unclear speech that means ask again.
The company also gives a figure: it measured Sparrow-2's conversation failure rate at 2.1 percent, nearly four times lower than the best alternative it tested. That measurement is Tavus's own, the systems compared are not named, and there is no independent verification.
The research note is on the Tavus blog.