buildOne report1 publisher Interfaze released interfaze-1-lite, an Apache 2.0 open-weight model for OCR and speech that returns confidence scores and bounding boxes with its answers. Teams can send doubtful rows to a reviewer once they have tested how the scores behave on their own documents.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence45
buildConfirmed7 publishers Microsoft's MAI-Transcribe-2 covers 60 languages with speaker labels and word timestamps, though streaming is not among its documented features. Live voice products get a fast model for replies and still need another way to hear the caller.
Perspective Coverage
8 publishers
- Builder
- Builder 53%
- Operator
- Operator 30%
- Investor
- Investor 17%
Reality
- Evidence35
- Adoption15
- Hype gap−35
- Incentives65
- Confidence70
buildOne report1 publisher David AI's 21-language benchmark found speech systems' error rates spread an average of 25.4 points in five Indic languages, against 3.7 in the tightest five. In those languages, choosing a system is a 25-point decision that an English score can make look like a close call.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives60
- Confidence50
Modulate raised $25 million to put models that analyze raw call audio, not transcripts, in front of more developers. Its pitch to teams running voice agents is a separate layer that spots cloned voices and grades how agents handle callers.
Perspective Coverage
3 publishers
- Builder
- Builder 40%
- Operator
- Operator 33%
- Investor
- Investor 27%
Reality
- Evidence58
- Adoption45
- Hype gap+20
- Incentives65
- Confidence60
buildConfirmed6 publishers Gemini 3.5 Transcribe ships as a sub-second streaming endpoint and a batch one, with filler-word removal and formatting inside the model. The catalog and the pricing page have not caught up.
Perspective Coverage
6 publishers
- Builder
- Builder 48%
- Operator
- Operator 36%
- Investor
- Investor 16%
Reality
- Evidence55
- Adoption35
- Hype gap+25
- Incentives70
- Confidence60
buildOne report1 publisher NVIDIA reports a 14.72% diarization error rate for the 100M-parameter Nemotron 3 Diarization, first on Voice Arena's initial leaderboard. Its Hugging Face preview page confines the weights to internal testing and evaluation.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+30
- Incentives65
- Confidence60
buildOne report1 publisher The author of a Crimean Tatar recogniser scored it for the first time at 34.6% word error, then found four of his audiobooks sitting in the data twice under filenames a name check had cleared.
Reality
- Evidence55
- Adoption12
- Hype gap−12
- Incentives30
- Confidence57
buildOne report1 publisher An RTX 3090 holding transcription and embedding models in VRAM averaged 25 W over the month at a cost of 2 euros, while the same post puts hardware amortisation at about 25 euros a month, twelve times the power bill.
Reality
- Evidence35
- Adoption10
- Hype gap+25
- Incentives50
- Confidence40
buildOne report1 publisher A developer's write-up of a medical transcription pipeline cites a 2025 study finding that under noise Whisper puts confidence above 0.7 on tokens that are wrong, and that the overconfidence grows as the signal-to-noise ratio falls.
Reality
- Evidence35
- Adoption10
- Hype gap+28
- Incentives65
- Confidence40
buildOne report1 publisher Apple's Kids Category rule and COPPA's definition of a child's voice both point a preschool voice product at local processing, and Whisper's error on children's speech falls by nearly a factor of three from tiny to large-v3.
Reality
- Evidence47
- Adoption38
- Hype gap+9
- Incentives71
- Confidence54
buildConfirmed2 publishers Granite Speech 5.0 Turbo CTC drops the language model and three quarters of its output steps. The speed looks real; the accuracy figures are IBM's own.
Reality
- Evidence63
- Adoption18
- Hype gap+16
- Incentives72
- Confidence68