Cactus Compute released Whistle, an open 16.9MB speech-to-text model that runs on a CPU in the same Needle engine as its tool-calling language models. The accuracy and speed figures are the company's own, so they need rerunning on real devices and audio.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives60
- Confidence45
Gemini 3.5 Transcribe ships as a sub-second streaming endpoint and a batch one, with filler-word removal and formatting inside the model. The catalog and the pricing page have not caught up.
Perspective Coverage
6 publishers
- Builder
- Builder 48%
- Operator
- Operator 36%
- Investor
- Investor 16%
Reality
- Evidence55
- Adoption35
- Hype gap+25
- Incentives70
- Confidence60
Mahidol's Biomedical and Data Lab publishes Thai Whisper fine-tunes that are free to use commercially, with 6.59 and 7.42 WER on Common Voice 13. Both were scored with the Deepcut tokenizer, and that shared segmentation is the reason the two numbers sit on one scale.
Reality
- Evidence45
- Adoption20
- Hype gap0
- Incentives30
- Confidence55
The early-bird rate on MAI-Transcribe-2 sits 72% below the line's April price, which saves a 100,000-hour buyer $26,000 a year and takes 26 cents out of every 36 a standalone vendor charged for the same hour.
Reality
- Evidence44
- Adoption21
- Hype gap+33
- Incentives74
- Confidence45
Muse Voice Transcribe's benchmark lead rests on three-tenths of a point across about eight hours of English audio, which is thinner evidence than the $0.18 hourly rate sitting underneath it.
Reality
- Evidence58
- Adoption28
- Hype gap+24
- Incentives68
- Confidence60
ArmBench-ASR v0.1 ranks nearly 30 systems on 20.7 hours of Armenian audio. The headline order flips on read speech, and every model degrades badly on movie dialogue.
Reality
- Evidence57
- Adoption22
- Hype gap+12
- Incentives58
- Confidence54