Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance
Oct 5, 2026, 11:44 PM · Hugging Face

Falcon-Emirati shows a small, deliberately trained model can speak a dialect that giant general models flatten, and that matters for every community AI currently talks past.
Why it matters
Modern Standard Arabic is the Arabic of news and textbooks, but in the UAE everyday life happens in Emirati Arabic, a Gulf dialect with its own vocabulary, idioms and poetry. The Falcon team has released Falcon-Emirati-7B, built on its Falcon-H1-Arabic family, to understand and answer in that dialect the way a native speaker would.
The headline result is about register, not trivia. On the team's open-ended tests, Falcon-Emirati scored 0.52 on dialect fidelity, meaning how often it actually answered in Emirati, while the competing models it tested scored between roughly 0.00 and 0.05. Those models often knew the answer but replied in formal Arabic anyway.
From the desk
We think this is one of the more quietly important kinds of AI work. A model that answers a grandmother's question in textbook Arabic isn't wrong exactly, but it isn't talking to her. The team's finding that scale alone doesn't buy dialect competence, with some of the largest multilingual models scoring below smaller dialect-aware ones, is a useful corrective to the idea that bigger always means better for everyone.
The method is honest about how hard this is. Emirati is mostly spoken, so there's little written text to learn from. The team combined dialect text crawled from Emirati sites and forums, formal-Arabic material about Emirati culture, and a large volume of synthetic dialect data constrained by glossaries and style rules. They describe much of the process as trial and error, which matches what anyone who's tried to adapt a model to a low-resource dialect would expect.
We'd weigh the evaluation carefully. The team built the Alyah benchmark with the community, where the model scores 84.83%, and the open-ended scoring used another company's model as judge. That's standard practice, but home benchmarks and AI judges both deserve outside checks. To its credit, the post flags the model's limits on rare expressions and asks for feedback from Emirati speakers.
The risk to name is synthetic data. Generating a lot of dialect text with a model, even with strict rules, can quietly standardize a living dialect around what the generator thinks it sounds like. If dialect models built this way scale, the communities they serve need a real voice in testing them. Done right, this is the blueprint for AI that meets people in their own language.
Context
Falcon-H1-Arabic pairs Mamba-style state space layers with transformer attention in each block, comes in 3B, 7B and 34B sizes, and was trained on a mix of formal and dialectal Arabic. The team chose the 7B size as the best balance of quality and cost. On a separate cultural-appropriateness test using 283 UAE scenarios, Falcon-Emirati scored 85.57%, ahead of the three other models tested.
Who feels it
- Emirati Arabic speakers
- A chatbot that answers in their own dialect, with cultural context, rather than defaulting to formal Arabic.
- Builders of regional and low-resource language AI
- A documented recipe showing that targeted data and evaluation beat raw scale for dialect work.
- Public services and businesses in the UAE
- Possible foundation for more natural local-language interfaces, with the team advising caution for sensitive or official uses.
What to watch
- Independent evaluations of Falcon-Emirati by native speakers outside the team
- Whether Falcon-Emirati and the Alyah benchmark see wider availability and community adoption beyond the team's chat platform
- Similar dialect-specialized models for other Arabic varieties