source-explainer
Falcon-Emirati's Dialect Score Is Real, Self-Graded, and Still Unverified
TII says Falcon-Emirati-7B scored 84.83% on its own new Emirati dialect benchmark, Alyah, but that number is self-graded, and no outside test yet confirms it.
Technology Innovation Institute's new Falcon-Emirati-7B model scored 84.83% on a benchmark called Alyah, built from 1,173 samples of greetings, etiquette, figurative language and poetry collected from native Emirati speakers. The Abu Dhabi lab says the model, fine-tuned from its Falcon-H1-Arabic base on authentic Emirati forum text, curated cultural material and glossary-constrained synthetic data, also posted a "dialect fidelity" score of 0.52 when prompted in Emirati, versus near zero for rival models. Those numbers come from TII's own blog post and its own benchmark, graded by TII's own human evaluators. That is not a flaw in the research so much as a gap in what has actually been tested: no independent group has yet replicated the result on Alyah, and the one outside study that measures the same problem tells a more complicated story.
That outside study is ArabCulture-Dialogue, a benchmark published in August by MBZUAI, a separate Abu Dhabi AI university, covering thirteen national Arabic dialects with native speakers tested in-country. Its headline finding, reported by Fast Company Middle East, is that large language models correctly recognize or select a culturally appropriate response roughly 95% of the time but their accuracy drops to around 50% when they have to generate dialect text themselves. The MBZUAI researchers single out Emirati and North African dialects as the hardest cases across the industry. TII's claim is specifically a generation claim, measured by dialect fidelity, which makes the comparison pointed: if 50% generation accuracy is the field's current ceiling on exactly this dialect, a self-reported 0.52 fidelity score looks less like a solved problem and more like a result sitting just above a known failure line, on a different test set, graded by a different set of judges.
None of this means Falcon-Emirati's result is wrong. Alyah and ArabCulture-Dialogue measure different things in different ways, and TII's own paper is explicit that standard NLP metrics cannot capture whether dialect text actually sounds right to a native ear, which is why it leaned on human judges rather than automatic scoring. That is a reasonable methodological choice. But it also means the 84.83% figure describes performance on a benchmark TII built, scored by evaluators TII selected, with no published detail in the blog post on how those native-speaker judges were recruited, how many dialectal registers or regions of the UAE they represented, or how disagreements between them were resolved. A score is only as informative as the test it comes from, and right now Alyah is not a shared instrument the field can check TII's number against. It is a private yardstick TII also built.
The release itself was bigger than one model. TII put out Falcon-Emirati alongside Falcon-ASR, a 1.6-billion-parameter multilingual speech-to-text model covering Emirati audio, and Falcon-OCR-Arabic, on October 6, as part of what the company is framing as a broader push to handle Arabic text, speech and vision together. Falcon-Emirati itself sits downstream of Falcon-H1 Arabic, the hybrid Mamba and Transformer model TII launched in January and marketed as topping the Open Arabic LLM Leaderboard while beating Llama-3.3-70B and Qwen2.5-72B despite running at a fraction of the parameter count. TII is the applied-research arm of Abu Dhabi's Advanced Technology Research Council, led by chief executive Dr. Najwa Aaraj, with the Falcon program run by chief researcher Dr. Hakim Hacid.
The dialect push also lands inside an active regional contest over whose Arabic model counts as sovereign. Jais, built by G42's Inception unit with MBZUAI, trained on a larger Arabic-specific corpus of 1.6 trillion tokens and has staked its claim on formal Modern Standard Arabic and regulated-sector use. TII is making the opposite bet, that dialectal fluency and latency are the differentiator, and the UAE is backing both horses through separate institutions in the same city. That duplication is itself a signal: Abu Dhabi has folded large language models into formal state strategy, having made its National AI System an advisory member of Cabinet in January, under the UAE AI Strategy 2031. Dialect fluency is not just a technical benchmark in that context, it is a claim about cultural infrastructure the state wants to own rather than license from a foreign hyperscaler.
There is also a recurring question about what "open" means in the Falcon name itself. Earlier Falcon releases drew criticism on the Hugging Face forums and the Open Source Initiative's discussion board for a license that imposed a 10% revenue royalty on commercial users above $1 million and restricted hosting providers, terms that OSI contributors argued made the "open source" framing misleading. TII has loosened the strictest versions of that license over successive releases, but the pattern means the license attached to Falcon-Emirati itself is worth reading before assuming the model is as open as the blog post implies, separate from whether its dialect scores hold up under outside testing.
What would actually settle the question raised by TII's own numbers is simple and has not happened yet: someone outside TII running Falcon-Emirati against ArabCulture-Dialogue, or running a rival model against Alyah, using judges neither lab selected. Until that happens, the honest reading of this release is narrower than the headline score suggests. A model has been built that TII says understands Emirati dialect at a level competitors do not. A test exists that TII itself designed to show that. What has not yet happened is anyone else checking the number against a yardstick they did not build themselves.