Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

TII's Falcon-Emirati-7B tops a dialect benchmark its own researchers helped build

TII's Falcon-Emirati-7B scores 84.83% on Alyah, an Emirati-dialect benchmark its own researchers helped develop. The model is small enough to trial cheaply for Gulf-facing chat, though every published score so far comes from TII's own testing.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying TII's Falcon-Emirati-7B tops a dialect benchmark its own researchers helped build
Generated illustration

What happened

  • Alyah is a multiple-choice test of 1,173 questions collected by hand from native Emirati speakers, covering everyday expressions, etiquette, figurative language, heritage and poetry.
  • In open-ended tests graded by Gemini 3.7 Flash, the model scored 0.52 for dialect fidelity, while the best rival, ALLaM-7B-Instruct-preview, scored 0.05.
  • It builds on TII's 7B Falcon-H1-Arabic from January, a hybrid that runs Mamba state-space components alongside Transformer attention in parallel blocks.
  • TII released the model on October 6, and users can try it through TII's Falcon chat platform.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Gulf-facing teams now have a 7B dialect model to trial against their current stack, and whether it ships should turn on their own native-speaker grading of real conversations.
  • constraint A multiple-choice knowledge score cannot stand in for a register test on generated replies, so the Alyah figure settles little about how the model answers actual users.
  • capability If the open-ended result holds up, answering in Emirati register does not require a large model, since the 7B beat two 27B models on dialect fidelity.
  • cost Checking the register claim properly needs native Emirati reviewers, the same input TII relied on in development, because the published comparison depends on an LLM judge.

Alyah tests recognition. Choosing among fixed options on etiquette or figurative language [8] is a different task from writing a reply in Gulf phrasing with nothing to choose from. RuntimeWire's write-up makes the same point: the score measures specialized dialect knowledge and does not by itself show how the model handles open-ended conversation or real applications [10]. The 84.83% works out to about 995 correct answers [15]. TII says that put the model ahead of the other Arabic and multilingual models it compared [9]. For that number to carry over to a production assistant, the assistant's users would need to be asking Alyah-style knowledge questions.

The open-ended test is closer to the product problem, and it is the better-designed half of the release. TII scored correctness separately from whether the answer was actually in Emirati Arabic [11]. A model that knows the answer but replies in formal Arabic still loses on dialect fidelity. Falcon-Emirati-7B's lead over the best rival on that score is about tenfold [16]. By TII's figures, Gemma 3 27B scored 0.03 and Jais-2-8B-Chat 0.02 [12]. Fanar-2-27B-Instruct scored effectively zero [12], a result with the virtue of being unambiguous. Both 27B models have nearly four times the parameters of TII's model [17].

The weak point is the grader. Gemini 3.7 Flash decided what counted as Emirati [11]. Emirati is mostly spoken, and it appears far less in online writing than MSA or even other Gulf and Levantine dialects [2]. The dialect-fidelity scores are only as reliable as the judge's grasp of a dialect with little written text to learn from. TII used native-speaker reviews during development to check whether responses sounded natural and culturally appropriate [4]. The published head-to-head rests on the LLM judge [11].

TII is candid that the training recipe was found by experiment. Its post says there is no documented playbook for how much dialect data is enough, how to mix it with MSA, or which stage among continued pre-training, SFT and preference optimization matters most [14]. A lot of the build "came down to trial and error," TII wrote [13]. The data came from three places: dialect text crawled from Emirati websites and forums, MSA writing about Emirati culture and identity, and synthetic dialect examples generated with glossaries and style rules [3].

The size argument is also TII's, and it is made inside TII's own model family. The 34B model "would likely push quality a bit further, but at a training and serving cost that doesn't make sense for a dialect-specialized chat model," TII wrote, while the 3B "doesn't leave enough headroom" [6]. One adoption cost has nothing to do with parameter count. The base model runs Mamba state-space layers in parallel with attention [5]. I'd confirm a serving stack handles those layers before comparing its cost with a plain Transformer of the same size.

The test I would run keeps TII's split and changes the grader. It would use real Gulf user turns, with replies from Falcon-Emirati-7B and from the incumbent model instructed to answer in Emirati dialect. Native speakers would score correctness and dialect use separately, as TII's open-ended test did [11].

What to watch

  • An evaluation of Falcon-Emirati-7B run by someone other than TII, on Alyah or on open-ended Emirati conversation.
  • Native-speaker scoring of TII's open-ended comparison, showing whether Gemini 3.7 Flash's dialect-fidelity grades agree with human judges.
  • Whether TII's release terms allow commercial self-hosting beyond the Falcon chat platform.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories