alyah-emirati-benchmark
收藏资源简介:
# Alyah ⭐️: Emirati Dialect Benchmark for Arabic LLMs  <div align="center"> 📝 [Blogpost](https://huggingface.co/blog/tiiuae/emirati-benchmarks) | 🔧 [Code on GitHub](https://github.com/huggingface/lighteval/pull/1117) </div> ## Dataset Summary **Alyah** (الياه in Emirati means North Star) is a manually curated evaluation benchmark for assessing the performance of Arabic large language models on the **Emirati dialect**. It is designed to measure not only linguistic competence, but also cultural awareness, pragmatic understanding, and sensitivity to social norms as expressed in everyday Emirati Arabic. The benchmark consists of **1,173 multiple-choice questions**, each with **four candidate answers** and **exactly one correct option**. Alyah targets scenarios where dialectal fluency and culturally grounded reasoning are required, going beyond Modern Standard Arabic (MSA). --- ## Motivation Despite Arabic being one of the most widely spoken languages globally, most existing benchmarks focus on Modern Standard Arabic, which differs substantially from how Arabic is used in daily life. Dialects—particularly Gulf and Emirati Arabic—remain underrepresented in evaluation settings, leading to blind spots in real-world model performance. Alyah was created to address this gap by providing a high-quality, dialect-focused benchmark grounded in authentic Emirati usage. The goal is to support the development and evaluation of models that better serve users in the UAE and contribute to more inclusive Arabic language technology. --- ## Dataset Composition All samples in Alyah were **collected manually from native Emirati speakers**, ensuring linguistic authenticity and cultural accuracy. For each question, large language models were used to **synthetically generate distractor choices**, which were then reviewed for plausibility and semantic proximity. To avoid positional bias, the **index of the correct answer is randomly distributed** across the four options. The dataset spans a diverse set of categories: | Category | Number of Samples | Difficulty | |---------------------------------|-------------------| -------------------| | Greetings & Daily Expressions | 61 | Easy | | Religious & Social Sensitivity | 78 | Medium | | Imagery & Figurative Meaning | 121 | Medium | | Etiquette & Values | 173 | Medium | | Poetry & Creative Expression | 32 | Difficult | | Historical & Heritage Knowledge| 89 | Difficult | | Language & Dialect | 619 | Difficult | This distribution emphasizes dialect-specific linguistic phenomena while maintaining strong coverage of culturally and socially grounded knowledge. --- ## Data Format Each sample follows a simple multiple-choice structure with the following fields: * `query`: Question text written in Emirati dialect or culturally grounded Arabic * `option_1`: first candidate answer * `option_2`: second candidate answer * `option_3`: third candidate answer * `option_4`: fourth candidate answer * `correct_answer`: Index of the correct answer (1–4) * `category`: One of the predefined semantic or cultural categories The dataset is suitable for **zero-shot and few-shot evaluation** of both base and instruction-tuned language models. --- ## Intended Use Alyah is intended **strictly as an evaluation benchmark**. It can be used for: * Comparing Arabic-native, multilingual, and regionally adapted models * Analyzing the effect of instruction tuning on dialectal understanding * Category-level analysis of strengths and weaknesses in Emirati Arabic It is **not intended for training** language models. --- ## Reference Evaluation As a reference, we evaluated **53 models** on Alyah, including **22 base models** and **31 instruction-tuned models**, spanning Arabic-native, multilingual, and adapted systems. Accuracy on the multiple-choice task is used as the primary metric. ### Base Models (Reference Results) | Model | Accuracy | |------|----------| | google/gemma-3-27b-pt | 74.68 | | tiiuae/Falcon-H1-34B-Base | 73.66 | | FreedomIntelligence/AceGPT-v2-32B | 67.35 | | google/gemma-3-4b-pt | 63.17 | | QCRI/Fanar-1-9B | 62.75 | | tiiuae/Falcon-H1-7B-Base | 60.78 | | meta-llama/Llama-3.1-8B | 58.23 | | Qwen/Qwen3-14B-Base | 57.29 | | inceptionai/jais-adapted-13b | 56.01 | | Qwen/Qwen2.5-32B | 53.03 | | FreedomIntelligence/AceGPT-13B | 50.81 | | Qwen/Qwen2.5-72B | 47.91 | | Qwen/Qwen2.5-14B | 46.8 | | google/gemma-2-2b | 41.86 | | tiiuae/Falcon3-7B-Base | 41.43 | | Qwen/Qwen3-8B-Base | 40.75 | | tiiuae/Falcon-H1-3B-Base | 40.41 | | Qwen/Qwen2.5-7B | 36.57 | | Qwen/Qwen2.5-3B | 35.29 | | meta-llama/Llama-3.2-3B | 35.12 | | inceptionai/jais-adapted-7b | 33.5 | | Qwen/Qwen3-4B-Base | 27.45 | ### Instruction-Tuned Models (Reference Results) | Model | Accuracy | |------|----------| | falcon-h1-arabic-7b-instruct | 82.18 | | humain-ai/ALLaM-7B-Instruct-preview | 77.24 | | google/gemma-3-27b-it | 74.68 | | falcon-h1-arabic-3b-instruct | 74.51 | | Qwen/Qwen2.5-72B-Instruct | 74.6 | | CohereForAI/aya-expanse-32b | 73.66 | | Navid-AI/Yehia-7B-preview | 73.32 | | FreedomIntelligence/AceGPT-v2-32B-Chat | 72.8 | | Qwen/Qwen2.5-32B-Instruct | 71.61 | | tiiuae/Falcon-H1-34B-Instruct | 71.1 | | meta-llama/Llama-3.3-70B-Instruct | 69.74 | | QCRI/Fanar-1-9B-Instruct | 69.22 | | tiiuae/Falcon-H1-7B-Instruct | 65.13 | | CohereForAI/c4ai-command-r7b-arabic-02-2025 | 64.54 | | silma-ai/SILMA-9B-Instruct-v1.0 | 63.94 | | FreedomIntelligence/AceGPT-v2-8B-Chat | 63.43 | | CohereLabs/aya-expanse-8b | 61.21 | | yasserrmd/kallamni-2.6b-v1 | 61.13 | | yasserrmd/kallamni-4b-v1 | 60.7 | | microsoft/Phi-4-mini-instruct | 58.57 | | tiiuae/Falcon-H1-3B-Instruct | 57.12 | | silma-ai/SILMA-Kashif-2B-Instruct-v1.0 | 48.51 | | Qwen/Qwen2.5-7B-Instruct | 45.44 | | google/gemma-3-4b-it | 46.12 | | meta-llama/Llama-3.1-8B-Instruct | 46.29 | | meta-llama/Llama-3.2-3B-Instruct | 39.64 | | yasserrmd/kallamni-1.2b-v1 | 37.77 | | Qwen/Qwen3-4B | 26.26 | | google/gemma-2-2b-it | 26.00 | | Qwen/Qwen3-14B | 26.00 | | Qwen/Qwen3-8B | 25.66 | These scores are provided for reference and reproducibility and should not be interpreted as universal rankings beyond the scope of Alyah. --- ## Evaluation Setup All reference results reported in this benchmark were obtained using **LightEval** with **Hugging Face Accelerate** for multi-GPU execution. The evaluation treats Alyah as a multiple-choice task and reports accuracy as the primary metric. We submitted a [PR](https://github.com/huggingface/lighteval/pull/1117) to Lighteval to officially support Alyah but while waiting for the PR to be reviewed and accepted, we share a [public fork](https://github.com/amztheorytii/lighteval_em.git) to allow the usage of the benchmark. The following command illustrates how to reproduce the scores above: ```bash git clone https://github.com/amztheorytii/lighteval_em.git cd lighteval_em pip install -e .[multilingual] pip install language_data accelerate launch --multi_gpu --num_processes=8 -m lighteval \ accelerate \ model_name=silma-ai/SILMA-Kashif-2B-Instruct-v1.0,batch_size=8 \ 'alyah' \ --output-dir results \ --load-tasks-multilingual \ --save-details ``` This setup allows consistent evaluation across base and instruction-tuned models, supports distributed inference, and logs per-sample details for further analysis. --- ## Limitations Alyah does not aim to exhaustively cover all Emirati dialectal variation. Language use varies across regions, generations, and social contexts within the UAE. While care was taken in reviewing synthetic distractors, subtle biases may still be present. --- ## Ethical Considerations The dataset was curated with attention to cultural, social, and religious sensitivity. Content reflects commonly accepted Emirati contexts and norms. Alyah is intended for **evaluation**, not for training models on sensitive content. --- ## License [Falcon LLM Licence](https://falconllm.tii.ae/falcon-terms-and-conditions.html) --- ## Citation ```bibtex @misc{emirati_dialect_benchmark_2026, title = {Alyah: An Emirati Dialect Benchmark for Evaluating Arabic Large Language Models}, author={Omar Alkaabi and Ahmed Alzubaidi and Hamza Alobeidli and Shaikha Alsuwaidi and Mohammed Alyafeai and Leen AlQadi and Basma El Amel Boussaha and Hakim Hacid}, year = {2026}, month = {january}, } ```



