Build1 publisher2 min readPublished
David AI's benchmark finds speech systems 25 points apart on Indic conversation
David AI's 21-language benchmark found speech systems' error rates spread an average of 25.4 points in five Indic languages, against 3.7 in the tightest five. In those languages, choosing a system is a 25-point decision that an English score can make look like a close call.
The Engineer · Build desk

What happened
- Microsoft AI's MAI-Transcribe-2 had the lowest transcription error rate in 18 of the 21 languages, the authors say, and ElevenLabs' Scribe v2 led the other three, including English.
- The test material comes from 147 hours of unscripted conversations between paired native speakers, each recorded on a separate synchronized channel and transcribed verbatim.
- NVIDIA Nemotron 3 Diarization posted the lowest overall diarization error rate at 18.8%, and David AI says every dedicated diarization model beat every combined system.
- David AI published the evaluation harness and scoring code on GitHub under the MIT License, while the dataset requires accepting a non-commercial data-use agreement.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision A vendor shortlist built on English results would favour the system that leads in three languages; a team serving Hindi or Tamil has to rank systems on its own language's results.
- constraint The diarization scores come from one speaker per microphone, a cue mono call recordings lack, so Nemotron's 18.8% is no forecast for single-channel audio.
- capability With the harness and scorer MIT-licensed, a team can apply the same scoring to its own recordings and reference transcripts, independent of the dataset's non-commercial terms.
Both spreads measure how far apart the 14 systems land within a single language on the public evaluation set [1][2]. In Bengali, Hindi, Marathi, Tamil and Telugu the average gap is about 6.9 times the gap in the five most consistent languages [1]. A choice worth under four points of error rate in one group is worth about 25 points, on average, in the Indic group [2][3]. According to Runtimewire's account, the report argues that an English score can make competing models look closer than they are in many other languages [16].
The reference transcripts are verbatim. Fillers, repetitions, false starts, interruptions and backchannels stay in, and every transcript went through two annotators and a language expert [6]. This is careful annotation work, and it is also a scoring decision. A system built to emit clean, readable text gets charged for each filler it drops, unless the scorer normalises disfluencies before comparing. Runtimewire's write-up does not describe that normalisation step. The scoring code is public [12], so its text-normalisation function is the first thing to read before trusting a row of the table.
Every system made its most diarization errors in the conversations with the most overlapping speech [11]. The hardest-overlap tier holds 15 recordings [8]. A ranking inside that tier rests on a small sample. Runtimewire's own caveat is that the results stay tied to this dataset and its recording conditions [17]. For any row to transfer, a team's traffic has to resemble the recording setup: two people talking unscripted, one channel each [5][7].
The downloadable slice is smaller than the benchmark. The dataset card lists 1,035 clips and 42.1 hours of two-channel audio from 682 speakers [14]. The source corpus is 147 hours [5]. The public audio is about 29% of the corpus by hours [2]. French, Italian and Korean are held back pending an internal privacy review [15], so outsiders can check at most 18 of the 21 languages from the public files [3].
YC partner Diana Hu described the founders' thesis as a need for a "Common Crawl for audio" [18]. The release contains identifiable voices and requires accepting a data-use agreement [13]. That is a sensible call for recordings of real people, and a narrower supply than the broad one Hu's phrase describes [18].
What to watch
- Whether David AI releases the French, Italian and Korean recordings after its privacy review, so outsiders can check all 21 languages.
- Whether the public scoring code normalises fillers and false starts before computing error rates, which would change how clean-output systems rank.
- Independent runs of the harness on single-channel call audio, to test whether dedicated diarization models keep their lead without per-speaker channels.