Skip to content

Topic

Automatic speech recognition

Technology that converts spoken audio into text, used in transcription, voice assistants and conversational AI, spanning many languages and accents.

Current stories

buildConfirmed7 publishers

Microsoft's MAI speech lineup documents real-time use only on the voice-output side

Microsoft's MAI-Transcribe-2 covers 60 languages with speaker labels and word timestamps, though streaming is not among its documented features. Live voice products get a fast model for replies and still need another way to hear the caller.

Perspective Coverage

8 publishers
Builder
Builder 53%
Operator
Operator 30%
Investor
Investor 17%

Reality

Evidence35
Adoption15
Hype gap−35
Incentives65
Confidence70
productConfirmed3 publishers

Modulate raises $25 million to sell voice-agent builders a separate audio-analysis layer

Modulate raised $25 million to put models that analyze raw call audio, not transcripts, in front of more developers. Its pitch to teams running voice agents is a separate layer that spots cloned voices and grades how agents handle callers.

Perspective Coverage

3 publishers
Builder
Builder 40%
Operator
Operator 33%
Investor
Investor 27%

Reality

Evidence58
Adoption45
Hype gap+20
Incentives65
Confidence60
buildConfirmed6 publishers

Google splits transcription in two, and quietly absorbs your cleanup layer

Gemini 3.5 Transcribe ships as a sub-second streaming endpoint and a batch one, with filler-word removal and formatting inside the model. The catalog and the pricing page have not caught up.

Perspective Coverage

6 publishers
Builder
Builder 48%
Operator
Operator 36%
Investor
Investor 16%

Reality

Evidence55
Adoption35
Hype gap+25
Incentives70
Confidence60