Microsoft's MAI-Transcribe-2 covers 60 languages with speaker labels and word timestamps, though streaming is not among its documented features. Live voice products get a fast model for replies and still need another way to hear the caller.
Perspective Coverage
8 publishers
- Builder
- Builder 53%
- Operator
- Operator 30%
- Investor
- Investor 17%
Reality
- Evidence35
- Adoption15
- Hype gap−35
- Incentives65
- Confidence70
LiveKit Agents now speaks in Microsoft's MAI voices, including MAI-Voice-2-Flash, through an official plugin that covers text-to-speech only. Teams no longer write their own adapter, but every deployment still needs an Azure Speech key and region.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives45
- Confidence50
ElevenLabs released Eleven v4 Turbo for voice agents, reporting medians of about 100 ms inference latency and 150 ms to first speech. Moving an existing agent takes more than a model ID swap, since older voice clones need retraining and SSML break tags no longer work.
Perspective Coverage
3 publishers
- Builder
- Builder 45%
- Operator
- Operator 33%
- Investor
- Investor 22%
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+25
- Incentives70
- Confidence60
Gemini 3.5 Transcribe ships as a sub-second streaming endpoint and a batch one, with filler-word removal and formatting inside the model. The catalog and the pricing page have not caught up.
Perspective Coverage
6 publishers
- Builder
- Builder 48%
- Operator
- Operator 36%
- Investor
- Investor 16%
Reality
- Evidence55
- Adoption35
- Hype gap+25
- Incentives70
- Confidence60
Preterview's developer cut mid-answer interruptions from 22% to 3.1% by confirming Deepgram's 300ms silence cutoff with a completeness check. The fix added 140ms at the median, against about 900ms for a longer timer, and threw away one drafted reply in five.
Reality
- Evidence45
- Adoption8
- Hype gap+15
- Incentives40
- Confidence50
SpaceXAI's new transcription model costs the same as the one it replaces and drops into existing code, so the migration work is re-validating accuracy claims that rest mostly on the company's own test sets.
Reality
- Evidence58
- Adoption40
- Hype gap+30
- Incentives65
- Confidence62
A dev.to walkthrough gives a Nova Sonic voice agent and an Amplify AI Kit chat agent one long-term memory in Amazon Bedrock AgentCore Memory, where every record is filed under an actorId and a sessionId. Both channels already had short-term context.
Reality
- Evidence58
- Adoption12
- Hype gap+14
- Incentives52
- Confidence58
On the platform behind taabi Nexus, the model's only output is a validated rule document, and the code that dials a driver's intercom sits behind a seven-day replay, a role check and one person's approval.
Reality
- Evidence58
- Adoption24
- Hype gap−12
- Incentives62
- Confidence55
Google's September 15 release pairs a fast speech-to-speech model with one that narrates its own reasoning aloud while tools run. Clients now have to read interactionStatus to know when a turn is actually over.
Reality
- Evidence30
- Adoption20
- Hype gap+12
- Incentives60
- Confidence35
Nabeel Baghoor shipped a lecture note-taker before he shipped voice agents, and reports that his accuracy wins on the note-taking side came from the recording path and a domain vocabulary bias list, with provider benchmarks a distant third.
Reality
- Evidence34
- Adoption18
- Hype gap+12
- Incentives35
- Confidence42
A dev.to writeup argues Whisper is a batch model retrofitted for streaming. If your response budget is under a second on mid-range phones, that is a build decision, not a tuning exercise.
Reality
- Evidence30
- Adoption12
- Hype gap+30
- Incentives85
- Confidence55