Block News International

@2026 Block News International. All Rights Reserved.

Blends Media
A Blends Media Group Production

UAE’s MBZUAI Introduces Benchmark for Culturally Aware Arabic AI

Arry Hashemi
Arry Hashemi
Aug. 17, 2026
MBZUAI MBZUAI researchers developed ArabCulture-Dialogue to test how AI models interpret everyday cultural situations and communicate across Modern Standard Arabic and 13 national dialects. (Image source: WAM)

Researchers at Abu Dhabi’s Mohamed bin Zayed University of Artificial Intelligence have developed a benchmark that examines whether artificial intelligence systems can understand Arab cultural situations and respond in the dialect people actually use. Called ArabCulture-Dialogue, the dataset covers Modern Standard Arabic and dialects associated with 13 Arab countries.

The work addresses a gap between formal Arabic evaluation and everyday communication. Modern Standard Arabic, or MSA, is widely used in education, news and formal writing, while daily conversations generally take place in regional dialects. Those dialects can differ in vocabulary, grammar and social conventions, creating challenges that may not appear when an AI system is tested only in MSA.

A model might, for example, recognize the meaning of a custom at an Emirati wedding but struggle to explain it naturally in Emirati Arabic. One scenario in the dataset involves a guest carrying an incense burner through a wedding gathering to perfume attendees with oud as a gesture of hospitality.

Building Conversations Around Local Culture

ArabCulture-Dialogue contains 6,942 dialogues: 3,471 in MSA and corresponding versions in national dialects. The conversations span 12 areas of daily life and 54 more specific subjects, including weddings, food, parenting, agriculture, arts and games. Of the full set, 2,780 dialogues focus on customs specific to individual countries.

The researchers adapted material from an earlier dataset called ArabCulture, which consisted of single-turn cultural questions. Initial conversational drafts were produced with a large language model before being revised and localized by 26 native Arabic speakers, two from each participating country. Annotators were instructed to write natural dialogue rather than translate each sentence literally, and the research team prohibited them from using AI tools during the human-editing stages.

Several rounds of review were used to reduce shortcuts that could allow a model to guess an answer without understanding the cultural context. Annotators adjusted answer choices when the correct option was noticeably longer or used a recognizable opening phrase. Independent reviewers also checked dialect consistency, cultural accuracy and whether the MSA and dialect versions retained the same meaning.

Stronger at Recognition Than Conversation

The benchmark evaluates models through three tasks. The first asks them to select a culturally appropriate response from three choices. The second tests translation between MSA and a specified dialect, while the third asks a model to continue a conversation in a named national dialect. Together, the tasks separate the ability to identify cultural information from the ability to express it naturally.

GPT-5 and Gemini 2.5 Pro recorded accuracy of approximately 94% to 95% on the multiple-choice cultural-reasoning test, depending on the language setting and whether geographic information was provided. Smaller Arabic-focused and multilingual models produced mixed results. Without geographic context, Jais 2 8B Chat scored 74% in MSA and 71% in dialect, while SILMA 9B Instruct recorded 78.3% and 71.6%, respectively.

Performance weakened more clearly when models had to generate dialect rather than choose an answer. The study found that systems were generally better at translating dialect into formal Arabic than converting MSA into a particular dialect. Even output that appeared fluent could be difficult to identify as belonging to the requested country, suggesting that grammatical quality alone did not guarantee dialect accuracy.

National Differences Remain Difficult

Country-level results showed that customs shared across the Arab world were generally easier for the models than practices associated with one location. North African dialects presented some of the largest difficulties, while Emirati conversations were also challenging. Providing the country and region in the prompt improved some scores, indicating that geographic guidance helped models retrieve relevant cultural information.

The findings also expose a limitation in treating national borders as neat linguistic boundaries. Each of the 13 participating countries is represented by one national dialect category, although speech can vary substantially between cities, regions and communities within the same country. The authors acknowledge this simplification and identify broader geographic and demographic coverage as an area for future research.

Lead researchers Fajri Koto and Muhammad Dehan said that the results show a gap between recognizing culturally appropriate language and producing authentic dialect. The evidence does not mean current models lack all relevant cultural knowledge. Instead, it shows that claims of supporting Arabic may obscure uneven performance across the language’s formal register, regional speech and country-specific social contexts.

A More Demanding Measure of Arabic AI

The research arrives as developers place greater emphasis on Arabic-language AI systems, including models designed or adapted for users in the Middle East and North Africa. Benchmark results can help researchers determine whether improvements reflect genuine dialect capability or gains limited to standardized Arabic tests.

ArabCulture-Dialogue also offers a framework for examining how AI systems handle culturally situated conversations rather than isolated questions. A customer-service assistant, educational platform or public-sector chatbot may encounter informal language, regional expressions and expectations about politeness that cannot be evaluated through formal translation alone.

The paper was published in the proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, held in July 2026. Its central finding is narrower than a claim that AI does not understand Arab culture: leading systems can often recognize an appropriate response, but consistently generating the right national dialect remains a separate and more difficult task.