Skip to content

Build1 publisher3 min readPublished

Armenian ASR leaderboard: closed models take the top eight, then lose the domains that matter

ArmBench-ASR v0.1 ranks nearly 30 systems on 20.7 hours of Armenian audio. The headline order flips on read speech, and every model degrades badly on movie dialogue.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Armenian ASR leaderboard: closed models take the top eight, then lose the domains that matter
Generated illustration

What happened

  • Hrant Davtyan, founder and CEO of Metric, released ArmBench-ASR on August 20, giving Armenian speech recognition a common test across read speech, poetry, movie dialogue and narrated news. The report's primary source is the Hugging Face Newsroom.
  • The initial v0.1 benchmark evaluates nearly 30 open-weight and closed systems using 10,113 audio clips totalling about 20.7 hours.
  • Google's Gemini 2.5 Pro placed first with a strict combined word error rate of 14.31%.
  • HiSpeech's Armenian conversational model followed at 16.81% and Gemini 2.5 Flash placed third at 17.55%; lower scores are better.
  • The eight best combined WER results came from closed systems, while NVIDIA's Armenian FastConformer was the highest-ranked open model in ninth place at 20.21%.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Metric, an Armenian research centre, published ArmBench-ASR on August 20, a shared evaluation of Armenian speech recognition across read speech, poetry, movie dialogue and narrated news, according to a report on runtimewire.com citing the Hugging Face Newsroom [1]. The v0.1 release scores nearly 30 open-weight and closed systems on 10,113 audio clips totalling about 20.7 hours [2], and the interesting part is not the ranking but how quickly the ranking stops holding [5][13].

On strict combined word error rate, Google's Gemini 2.5 Pro placed first at 14.31% [3], HiSpeech's Armenian conversational model second at 16.81% and Gemini 2.5 Flash third at 17.55%, with lower being better [4]. Closed systems took all eight top places; NVIDIA's Armenian FastConformer was the best open model in ninth at 20.21% [5], a gap of 5.90 points from first [6]. That is the number a procurement deck will screenshot.

Then read the leaderboard by dataset. On Common Voice 26, FastConformer led at 7.70% WER, ahead of Gemini 2.5 Pro's 9.69% [13] - 1.99 points the other way [14]. HiSpeech's two Armenian models beat Gemini on strict WER for poetry, while Gemini led once text was normalized [15]. Metric's own conclusion is that model choice depends on the audio and on the output conventions you want [15].

The scoring mode moves the numbers as much as the model does. Strict scoring preserves punctuation and capitalisation; normalized scoring lowercases and strips punctuation after applying consistent Unicode, spacing and character rules to both references and model output [8]. Gemini 2.5 Pro's normalized combined WER was 6.42%, less than half its strict score [7]: 7.89 points, or about 55% of the strict figure, sits in formatting rather than in words [9]. If your pipeline post-processes text anyway, the strict column is charging you for errors you will never see.

The hardest domain is the one nobody benchmarks. Every evaluated model recorded its highest strict WER on Metric's private Movies collection, where the median reached 61.87%, against 17.29% on Common Voice 26 and 17.02% on FLEURS [10] - roughly 3.6 times the read-speech median [11]. Metric offers background noise as one possible factor, while noting that conversational delivery and produced audio differ from read speech in other ways too [12]. For anyone transcribing calls, meetings or media, the clean-read figure is not a conservative estimate; it is the wrong order of magnitude.

The construction is worth noting. Two components are public: Common Voice 26 and Google's FLEURS Armenian test split, for which Metric published a small set of transcript corrections after manual review [16]. Three are private: Poems (expressive literary readings over background music), Movies (conversational dialogue with produced-media noise) and Infocom (Armenian news text recited by a single speaker), all manually reviewed and corrected [17]. The private material is what makes the benchmark useful and also what outsiders cannot inspect, since it is not part of the release [22]. The average clip runs about 7.4 seconds [23], so these are short-utterance results, not long-form ones.

Metric says this Armenian-language work is a self-funded track in which its researchers spend their own time and budget [20]; Hrant Davtyan, credited as founder and CEO, is also an assistant professor at the American University of Armenia and a Metric co-founder, with Alexander Shahramanyan and Mariam Avetisyan among the credited researchers [21].

Watch whether v0.2 adds long-form or multi-speaker audio, whether the private sets stay private, and whether vendors start quoting normalized numbers without saying so [8][18][22].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories