Skip to content

Topic

Voice agents and live captioning

Interactive voice products - agents, real-time captions, live meeting transcription - where perceived response time governs usability.

Current stories

build7 publishers

Microsoft's MAI speech lineup documents real-time use only on the voice-output side

Microsoft's MAI-Transcribe-2 covers 60 languages with speaker labels and word timestamps, though streaming is not among its documented features. Live voice products get a fast model for replies and still need another way to hear the caller.

Perspective Coverage

8 publishers
Builder
Builder 53%
Operator
Operator 30%
Investor
Investor 17%

Reality

Evidence35
Adoption15
Hype gap−35
Incentives65
Confidence70
build3 publishers

ElevenLabs aims v4 Turbo at live voice agents with a self-measured 150 ms to first speech

ElevenLabs released Eleven v4 Turbo for voice agents, reporting medians of about 100 ms inference latency and 150 ms to first speech. Moving an existing agent takes more than a model ID swap, since older voice clones need retraining and SSML break tags no longer work.

Perspective Coverage

3 publishers
Builder
Builder 45%
Operator
Operator 33%
Investor
Investor 22%

Reality

Evidence50
Adoption
Insufficient
Hype gap+25
Incentives70
Confidence60
build6 publishers

Google splits transcription in two, and quietly absorbs your cleanup layer

Gemini 3.5 Transcribe ships as a sub-second streaming endpoint and a batch one, with filler-word removal and formatting inside the model. The catalog and the pricing page have not caught up.

Perspective Coverage

6 publishers
Builder
Builder 48%
Operator
Operator 36%
Investor
Investor 16%

Reality

Evidence55
Adoption35
Hype gap+25
Incentives70
Confidence60