← Back to Analysis

The Rise of Frontier Arabic AI

Gulf Arabic AI models, showing their foundation-model relationships, launch years, and reported sizes.
Visual by Jesse Marks · View full size

I have been digging deep into some of the emerging trends in Gulf AI, specifically around Arabic model development. This paper reflects some of this research and aims to make a wide range of political science and technical AI research on Arabic models accessible to a wider audience. For more technically informed readers, please feel free to reach out, comment, or correct where I may be miss a few things.

Both Saudi Arabia and the UAE are developing Arabic AI models that are rapidly approaching frontier capabilities and can reason across a wide range of Arabic dialects. If those models become genuinely competitive at the frontier, Gulf states could begin to accumulate advantages that do not depend on matching the United States or China across every layer of the AI stack.

The more important question is whether Arabic reasoning itself can become a durable source of technological capability, especially as evaluation, safety, and model development around Arabic remain less mature than they are in English.

The Politics of Building a Frontier Model

Saudi Arabia’s HUMAIN surprised many observers last week with the release of humain-m3, a 428-billion-parameter mixture-of-experts model developed on the MiniMax-M3 lineage and further pretrained on more than one trillion tokens of Arabic-native content. HUMAIN presents M3 as a frontier Arabic model, though that claim will require further independent testing in the coming weeks. Its choice of MiniMax also marks a departure from earlier Gulf Arabic model development.

Saudi Arabia’s ALLaM was built on Meta’s Llama 2, while major Emirati efforts such as Jais and Falcon were developed more independently inside the UAE.

The Saudi shift from Meta to MiniMax also fits HUMAIN’s broader approach to offering a wider model choice. CEO Tareq Amin has described the company’s strategy as building access to the “best models from the US, China, Europe, and the open-source ecosystem,” rather than locking HUMAIN into a single supplier or national technology stack.

https://www.linkedin.com/posts/tareq-amin_sajid-nahvi-i-agree-that-the-orchestration-share-7483539607062507521-X5RF/

This frames the shift from Llama 2 for ALLaM to MiniMax for M3 as a pragmatic decision to use whichever foundation model gives HUMAIN the strongest base for the system it wants to build. This is a pragmatic decision rather than any form of geopolitical pivot toward Chinese AI models.

Gulf Arabic AI models, showing their foundation-model relationships, launch years, and reported sizes.
Author produced visual map of Arabic AI models © Jesse Marks · View full size

M3 is useful because it shows the political complexity of building frontier Arabic AI. Gulf states have spent heavily to gain greater control over compute, data centers, energy, and talent, but much of that infrastructure remains tied to American hardware and the broader U.S. AI ecosystem. The model layer is less straightforward. Access to American compute does not mean that an American model will always be the best foundation for a specialized Arabic system.

This also comes as Washington seeks to limit the exposure of American-backed AI infrastructure to Chinese technology and models.

That creates a difficult choice for Gulf developers like HUMAIN. Chinese models have become increasingly capable, relatively inexpensive, and easier to adapt, but U.S. authorities accuse Chinese developers of using large-scale distillation to improve Chinese models, such as MiniMax, Kimi, and DeepSeek. HUMAIN M3 builds on one of these models with Arabic data and post-training inside an AI ecosystem that remains heavily dependent on American compute.

That is perhaps the most noteworthy case of a hybrid stack, but much of the energy, data center infrastructure, deployment environment, and Arabic-specific development is located in Saudi Arabia. The interesting question is when does M3 stop being a “Chinese open model” and become a “Saudi model”? This is a debate we may see play out in the near future.

The UAE, meanwhile, is pursuing an indigenous path through Falcon and its wider research ecosystem. Abu Dhabi’s Technology Innovation Institute has trained Falcon-H1 Arabic on native Arabic data across Modern Standard Arabic and several regional dialects. While Saudi Arabia has so far relied more on foreign foundation models like Meta and MiniMax and localized around them, the UAE has moved further toward developing its own model families and integrating its own Arabic data.

Does Frontier Arabic AI Matter

Most leading frontier models can already produce competent Arabic across a wide range of dialects. They cannot, however, reason in Arabic. Saudi and Emirati models may eventually achieve stronger reasoning performance in Arabic, particularly in areas where language, dialect, and regional context shape answer quality. That includes legal interpretation, government administration, Islamic finance, social context, regional history, and movement between Modern Standard Arabic and everyday dialect.

A model that performs materially better in these areas would give Gulf states a capability that foreign frontier labs may not be able to reproduce simply by translating an English-centered model into Arabic. It also is one that is more approachable to native Arabic speakers across wide range of contexts.

A recent Saudi dialect benchmark tested Claude Opus 5, Gemini 3.7, GPT-5.6, and Kimi K3 on expert-authored prompts involving idioms, pragmatics, vocabulary, and culturally embedded meaning. None exceeded 55 percent overall, and many of the errors came from misunderstanding what speakers meant in context rather than from simply producing false information.

The results point to a gap between producing fluent Arabic and reliably interpreting what Saudi speakers mean when meaning depends on dialect, register, or context.

Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, and Maxim Legg, “Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models,” arXiv preprint arXiv:2608.29990 (2026), https://arxiv.org/abs/2608.29990.

This gap gives Gulf states room to build greater domestic control at the model layer. The performance of their respective Arabic models can be improved with a range of inputs that are harder to obtain from outside the region,including high-quality Gulf dialect data, government and legal terminology, technical language used across specific sectors, and feedback from Arabic users.

These inputs become more valuable when paired with post-training, evaluation, and engineers who understand how Arabic is actually used in government, business, and society. Much of that local language, institutional knowledge, and deployment experience is difficult for foreign labs to reproduce at the same depth.

The longer-term value of the Saudi approach will depend on whether HUMAIN can carry what it learns from M3 into future models. Saudi Arabia may not control the underlying MiniMax foundation, but it can still build its own Arabic data, post-training methods, benchmarks, safety tools, and engineering expertise around it.

If those capabilities can be reused on a different foundation model later, HUMAIN becomes less dependent on any one foreign model provider and gains more control over the Arabic layer of the stack.

This becomes particularly important if the United States restricts the use of Chinese-origin models on American-backed AI infrastructure, or if HUMAIN eventually shifts to another foundation model. The sovereignty test is whether the Arabic capability built around M3 can move with HUMAIN across different model foundations rather than remaining tied to MiniMax.

Evaluating an Arabic Frontier Model

Measuring Arabic reasoning models using translated English benchmarks will be difficult. Translations that lack linguistic context can change the difficulty of a question, remove ambiguity, or miss critical references or forms of reasoning that are embedded within the language itself. Arabic also varies widely across Modern Standard Arabic and the breadth of regional and subregional dialects, which can, at times, be mutually unintelligible.

In practice, this means model performance judged on Modern Standard Arabic may say relatively little about how a model performs in the language many users actually use.

A credible reasoning benchmark would need to be designed in Arabic by native Arabic speakers with questions whose meaning, context, and difficulty flow naturally from Arabic rather than from machine-translated English prompts. These evaluations could draw on legal interpretation, cultural references, government terminology, and other areas where context shapes meaning.

Building these types of evaluations requires researchers who understand model evaluation methods, have a native, fluent grasp of Arabic, and can combine these to develop novel methods for testing Arabic AI models.

Saudi and Emirati institutions are already building substantial Arabic evaluation infrastructure. Abu Dhabi’s Technology Innovation Institute runs QIMMA, a quality-controlled Arabic LLM leaderboard that evaluates dozens of models across more than 52,000 samples spanning cultural, legal, medical, safety, STEM, and coding tasks. Uniquely, a Chinese-built Qwen model outscored the other Arabic LLMs, including the UAE’s JAIS and Egypt’s Karnak, which is also built on Qwen.

Saudi-focused researchers have also developed new dialect benchmarks, including a recent Saudi Arabic evaluation built around expert-authored prompts that test idioms, pragmatics, vocabulary, and culturally embedded meaning. Broader regional efforts such as DialectalArabicMMLU now test models across Saudi, Emirati, Egyptian, Syrian, and Moroccan Arabic.

Together, these efforts are concentrating more Arabic-language data, dialect expertise, and model-testing capacity inside the region.

Leen AlQadi, Ahmed Alzubaidi, Mohammed Alyafeai, Maitha Alhammadi, Shaikha Alsuwaidi, Omar Saif Alkaabi, Basma Boussaha, and Hakim Hacid, “QIMMA قِمّة: A Quality-First Arabic LLM Leaderboard,” Hugging Face, April 21, 2026, https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard.

If Gulf-developed benchmarks gain wider use, they could also give governments and companies across the Arab world a more relevant basis for comparing models for local deployment. That would make it easier for Gulf-developed models to demonstrate advantages in areas where general-purpose frontier models still perform unevenly, particularly dialect, context, and region-specific reasoning.

The Emerging Arabic AI Safety Gap

If Saudi and Emirati models are approaching the frontier, as they claim, this signals an urgent need for AI safety experts to focus on the risks of frontier Arabic AI deployed in the region. So far, most AI safety organizations remain concentrated around frontier labs in the West. Recent releases of advanced models such as Anthropic’s Mythos and Fable have sharpened concern over how quickly frontier systems are gaining capabilities in cyber operations and other high-risk domains.

Leading American models are also being used by non-state actors and terrorist groups for malign purposes. Anthropic’s September 2026 threat intelligence report documented a Yemen-based cell, likely linked to the Houthis, using Claude Code to support missile-guidance and weapons work. Iranian state-aligned actors also used Claude for surveillance, target identification, and domestic monitoring. This raises important considerations for future Gulf AI safety.

First, as Arabic frontier models become more capable, the risk of exploitation by malign users is likely to increase. That makes safeguard robustness a first-order problem for Gulf developers. Arabic creates additional safeguard workarounds through dialect, code-switching, indirect phrasing, and colloquialism that may not trigger the same safety behavior as English or Modern Standard Arabic.

In July, AI researchers showed that relatively small shifts from Modern Standard Arabic into dialect could manipulate the models response and enable harmful prompts. Malicious users could exploit those differences to conceal intent, confuse the model, or weaken its safeguards.

Arabic-specific safety benchmarks are beginning to expose these gaps. SalamaBench found in its evaluation of Arabic-language models using more than 8,000 prompts across a range of safety categories that there is substantial variation in how models handle harmful requests. ASAS, a separate human red-teaming benchmark, also found weaknesses across leading Arabic-capable models. This suggests that AI safety behavior does not transfer consistently across languages.

ASAS researchers also found that non-human, model-based safety evaluation techniques were less reliable than human Arabic-speaking assessment. This means that policing frontier Arabic AI will require skilled human evaluators who can understand dialect, context, and implied intent, rather than using AI systems to police other AI systems in Arabic.

The second challenge is a significant institutional gap in AI safety and governance expertise in the Gulf region needed to address these future challenges. Frontier Arabic models will require a more robust local safety and governance ecosystem. Riyadh and Abu Dhabi will need more Arabic-speaking red teams, alignment researchers, computational linguists, benchmark designers, model auditors, and independent evaluators who can test systems for real-world misuse patterns.

That requires a wide range of actors, including state, industry, university, and third-party evaluators, to invest in cultivating that intellectual capability in Arabic. The robust corpus of AI safety and governance research mostly exists in English, and the institutions dedicated to preventing frontier safety risks conduct this work in English and European languages.

Bridging this gap will require a significant focus on translating that knowledge into Arabic and equipping AI researchers with the evaluation frameworks and skills to conduct similar research on frontier Arabic AI safety.

Arabic AI as a Regional Entrepôt

If Arabic AI continues to improve, Saudi Arabia and the UAE could become an important deployment layer for a group of Arab states. In that scenario, Gulf-based AI labs could provide governments and firms across the region with Arabic-capable models adapted to local dialects and specifications.

Wider deployment would create a data feedback loop that generates new Arabic-language interaction data, error cases, preference data, and sector-specific examples that can be used for future post-training and evaluation. That could give Gulf-developed Arabic AI a cumulative regional advantage over imported models.

As deployment expands, Gulf institutions would accumulate more Arabic data and practical experience across different markets, which could improve subsequent models and make them more attractive to additional users across the region.

Model deployment would also complement Gulf ambitions to export their AI stack to other countries in the region. For example, the UAE’s G42 partnership with Benya Technologies in Egypt aims to deliver Cairo AI and critical digital infrastructure, including data centers, telecommunications towers, and cloud technology.

If Gulf states can combine regional compute, cloud infrastructure, and Arabic models, they could become important intermediaries between global frontier developers and Arab markets. This would give Riyadh and Abu Dhabi influence over both the physical infrastructure and the model-deployment layer of AI adoption across the Arab world, while also expanding the regional data and deployment base available to their own AI systems.

Note on AI Usage: This paper used GPT Sol for copy editing and some visualizations.

Originally published in Coffee in the Desert on September 15, 2026.

The Rise of Frontier Arabic AI | The Gulf AI Monitor