Build1 publisherNot yet confirmed elsewhere3 min readPublished
TII says its 1.6B Falcon-ASR beats the best published Arabic word error rate by 2.25 points
TII's Falcon-ASR averaged a 20.92% word error rate on six Arabic test sets, under the 23.17% best published result. Its lead on Emirati speech, the dialect it was built around, comes from a test set TII assembled and scored itself.
The Engineer · Build desk

What happened
- On TII's internal Emirati evaluation, Falcon-ASR recorded 22.73% WER and 10.19% CER, the lowest of the systems TII chose to compare.
- One set of weights transcribes Arabic, English, French, Spanish and Portuguese without a language flag, returning text in whichever language was spoken.
- Transcripts can carry word-level timestamps that link each word to its position in the audio.
- The only announced way to try the model is a Hugging Face demo, and TII says API access and native applications are planned.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision A team shipping to Emirati speakers has to score candidates on the public Casablanca UAE subset or on its own recordings before counting the 4.07-point dialect margin as its own.
- cost On inference cost the dialect win is over a model running under twice Falcon-ASR's parameter count per token, so the efficiency gain a team can bank is far smaller than the 30-billion total implies.
- capability Subtitling and search features can take word positions straight from the transcript, and a multilingual pipeline needs no language-selection step before transcription.
- constraint Until the planned API ships, teams can spot-check recordings in a demo but cannot run Falcon-ASR over a production-sized sample of their own audio to see whether 20.92% transfers.
The 20.92% figure comes from a published protocol. The Open Universal Arabic ASR Leaderboard, maintained by the ELM Research Center, ranks systems by the equal-weight average WER across six test sets and reports CER beside it [7]. TII ran Falcon-ASR on the leaderboard's pinned manifests and compared the result with competitors' published averages as checked on 30 September 2026 [9]. So Falcon-ASR's score is TII's own run, and the competitors' scores are the leaderboard's [9]. The 23.17% it beat belongs to Audar-ASR-V1-Turbo [6]. A 2.25-point lead [5] is about 9.7% fewer word errors in relative terms [22]. TII puts the matching average CER at 8.79% [20].
Equal weighting decides how far that number travels. Each of the six sets counts the same in the average [7]. A model can therefore lead the average while trailing on the one set that sounds most like a given product's traffic. For 20.92% to hold in production, the production audio has to resemble the six-set mix. TII wrote in the announcement: "A model that handles a formal news broadcast may still struggle with a conversation in Emirati or with speech recorded over a phone line." [17]
The Emirati result has a different provenance. TII built that evaluation from held-out Emirati and Gulf recordings with human-validated transcripts [1]. Qwen3-Omni came second at 26.80% WER [2], 4.07 points behind [3], a relative gap of about 15% [23]. RuntimeWire wrote that the gap "remains a result from TII's own evaluation and the systems TII chose to compare" [8]. A public cross-check exists. The Casablanca evaluation data includes a UAE subset [21].
The size comparison needs a footnote. RuntimeWire lists Qwen3-Omni at 30 billion parameters with 3 billion active [2]. Against Falcon-ASR's 1.6 billion [4], it has roughly 19 times the total but under twice the active count [24].
I think the training recipe is aimed at the right audio. TII says it trained on Emirati, MSA, other Gulf and Arabic dialects, and English, with the aim of transcribing everyday speech including switches between languages [14]. It added background noise, overlapping speech, music, room reverberation, telephony effects, and speed and pitch changes, and applied the same treatment to the Emirati recordings to cover meetings and calls [15]. On English, it reports a 5.74% mean WER across the seven public sets used by the Hugging Face Open ASR Leaderboard [13]. The model extends TII's Falcon3-Audio work [16]. Two of the four credited authors, Abdul Muneer and Ludovick Lepauloux, co-authored that earlier research [18].
Beyond the demo and the planned API [11], the announcement does not mention downloadable weights or a license. RuntimeWire called the demo "the announced route to test it rather than evidence of a production service" [19].
What to watch
- A Falcon-ASR entry on the Open Universal Arabic ASR Leaderboard to sit beside TII's self-run 20.92%.
- Independent scores on the public Casablanca UAE subset comparing Falcon-ASR with Qwen3-Omni.
- TII's planned API terms, and whether downloadable weights and a license accompany it.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence40
- Adoption8
- Hype gap+20
- Incentives75
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
TII's internal evaluation of additional Emirati and Gulf speech uses held-out recordings and human-validated transcripts; on it Falcon-ASR achieved 22.73% WER and 10.19% CER, the lowest WER and CER among the systems compared.
ReportedSupportedSource: TII internal evaluation3 sources— create a free account to open themView cited source - [2]
In TII's internal Emirati comparison, the next-lowest WER was 26.80% for Qwen3-Omni, a model listed at 30 billion parameters with 3 billion active.
- [3]
Falcon-ASR's Emirati WER is 4.07 percentage points below Qwen3-Omni, the next best result.
- [4]
TII introduced Falcon-ASR, a 1.6 billion parameter speech recognition model for Arabic with a particular focus on the Emirati dialect, developed at the Technology Innovation Institute in Abu Dhabi; it also supports English, French, Spanish and Portuguese.
ReportedSupportedSource: TII announcement on Hugging Face blog2 sources— create a free account to open themView cited source - [5]
In TII's evaluation, Falcon-ASR achieved an average word error rate of 20.92% across six Arabic test sets, compared with the best published result of 23.17% in the leaderboard snapshot used, 2.25 percentage points better.
- [6]
The 23.17% best published result in the leaderboard snapshot TII used belongs to Audar-ASR-V1-Turbo.
- [7]
The Open Universal Arabic ASR Leaderboard, maintained by the ELM Research Center, ranks systems by the equal-weight average WER across six test sets and also reports character error rate; TII's Falcon-ASR evaluation follows this protocol.
- [8]
remains a result from TII's own evaluation and the systems TII chose to compare
ReportedSupportedSource: RuntimeWire, on the Emirati WER gap3 sources— create a free account to open themView cited source - [9]
TII evaluated Falcon-ASR on the same six benchmarks using the leaderboard's pinned manifests; competitor figures are the published leaderboard averages checked on 30 September 2026.
- [10]
Falcon-ASR supports word-level timestamps for transcriptions, linking each transcribed word to its position in the audio.
- [11]
Falcon-ASR can be tried in a Hugging Face Demo Space; API access and native applications are planned.
- [12]
All five supported languages use the same model weights without requiring a language flag; the output is a transcript in the language spoken.
- [13]
On the seven public English test sets used by the Hugging Face Open ASR Leaderboard, Falcon-ASR achieved a mean WER of 5.74%.
- [14]
TII trained Falcon-ASR on Emirati, Modern Standard Arabic, other Gulf and Arabic dialects, and English, aiming to transcribe everyday speech including dialectal forms and changes between languages.
- [15]
Training included background noise, overlapping speech, music, room reverberation and telephony effects, plus variations in speed and pitch, with the same treatment applied to Emirati recordings to cover conditions in meetings, calls and other everyday recordings.
- [16]
Falcon-ASR builds on TII's Falcon3-Audio work.
- [17]
A model that handles a formal news broadcast may still struggle with a conversation in Emirati or with speech recorded over a phone line.
ReportedSupportedSource: TII announcement text2 sources— create a free account to open themView cited source - [18]
The work is credited to Abdul Muneer, Ludovick Lepauloux, Rishabh Saraf and Shamsa Hamad; Muneer and Lepauloux co-authored TII's earlier Falcon3-Audio research.
- [19]
the announced route to test it rather than evidence of a production service
ReportedSupportedSource: RuntimeWire, on the Hugging Face demo2 sources— create a free account to open themView cited source - [20]
Across six Arabic test sets, TII reports an average character error rate of 8.79% for Falcon-ASR.
- [21]
Public evaluation data already includes Emirati: Casablanca has a UAE subset.
- [22]
Falcon-ASR's 2.25-point lead over 23.17% is about a 9.7% relative reduction in word error rate.
- [23]
Falcon-ASR's 4.07-point Emirati lead over Qwen3-Omni's 26.80% is about a 15% relative reduction in WER.
- [24]
Qwen3-Omni has roughly 19 times Falcon-ASR's parameter count in total but under twice it in active parameters.
Sources
1 independent publisher whose own reporting we read for this story.
- huggingface.coIntroducing Falcon ASR
1 article · October 7, 2026
- runtimewire.comTII says its 1.6B Falcon-ASR beats larger models on its Emirati test
1 article · October 7, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Arabic language AIFollow
- Automatic Speech RecognitionFollow
- AI BenchmarksFollow