Researchers from the Technology Innovation Institute have released Falcon-Emirati-7B, an open language model designed to understand and generate Emirati Arabic rather than relying solely on Modern Standard Arabic. Standard Arabic serves as the common written standard for news and textbooks across the region, but everyday conversation, humor, storytelling, and local nabati poetry rely on dialectal phrasing and cultural context that literal machine translation frequently misses. Building a dedicated dialect model addresses this gap by training an artificial intelligence system directly on the vocabulary, etiquette, and figurative expressions used in the United Arab Emirates.
The system is built on top of the Falcon-H1-Arabic model family, which uses a hybrid architecture that pairs State Space Models (Mamba) with Transformer attention inside each block. Combining both designs offers linear-time efficiency on long context sequences through Mamba while maintaining attention mechanisms to track long-range linguistic dependencies. While the Falcon-H1-Arabic base models span 3-billion, 7-billion, and 34-billion parameter scales with context windows extending up to 128,000 and 256,000 tokens, the team chose the 7-billion parameter size for Falcon-Emirati-7B. According to the developers, the 7B size balances training and inference costs against the capacity needed to retain cultural nuance, avoiding the higher serving expenses of the 34B version and the limited depth of the 3B variant.
Developing the model presented unique data challenges because Emirati is primarily a spoken dialect with limited written text available online compared to Standard Arabic or other regional dialects. To assemble training material, the team built a pipeline using three separate inputs: authentic web content curated from Emirati websites and forums, Modern Standard Arabic materials focused on Emirati heritage, history, and social norms, and synthetic data. To prevent synthetic generations from mimicking generic Gulf phrasing, the authors constrained the generation systems using dictionaries, glossaries, and strict rules tailored to local grammar and vocabulary.
To determine how to apply the data, the developers tested different training mixtures, adaptation stages, and supervision strategies across continued pretraining, supervised fine-tuning, and preference optimization. Native Emirati speakers conducted manual reviews throughout the process to evaluate conversational tone, cultural appropriateness, and naturalness. Quantitative tracking relied on Alyah, a multiple-choice benchmark containing 1,173 samples gathered directly from native speakers covering greetings, etiquette, poetry, and heritage knowledge.
On the Alyah benchmark, Falcon-Emirati-7B achieved an accuracy score of 84.83%, outperforming other evaluated Arabic and multilingual instruction-tuned models, including systems with much larger parameter counts. Hugging Face reports that the testing demonstrates scale alone cannot compensate for targeted dialect adaptation, as generalized multilingual systems struggled on localized cultural tasks.



