AraLingBench
收藏资源简介:
AraLingBench是一个包含150个问题、由专家编写的基准测试,专门隔离了阿拉伯语能力的五个基本支柱——语法、形态学、拼写、阅读理解和句法,用于诊断大型语言模型是否具备真正的语言能力而非表面流利度。该数据集包含由训练有素的阿拉伯语语言学家审查的原创项目,涵盖不同难度级别,采用零样本评估协议。
AraLingBench is an expert-developed benchmark containing 150 questions, which is specifically constructed to isolate the five core pillars of Arabic language proficiency: grammar, morphology, spelling, reading comprehension, and syntax, aiming to diagnose whether large language models (LLMs) possess genuine linguistic competence rather than just superficial fluency. This dataset includes original items reviewed by trained Arabic linguists, covers a range of difficulty levels, and adopts a zero-shot evaluation protocol.
AraLingBench 数据集概述
数据集简介
AraLingBench 是一个人工标注的基准测试,专门用于压力测试大型语言模型的阿拉伯语语言核心能力。该基准测试包含150个专家编写的问题,重点评估阿拉伯语的语言结构理解能力。
核心特征
基准规模
- 问题总数:150个多项选择题
- 类别数量:5个语言学类别
- 难度分布:33%简单、49%中等、17%困难
语言学类别
每个类别包含30个题目:
- 语法(Grammar)
- 形态学(Morphology)
- 拼写与正字法(Spelling & Orthography)
- 阅读理解(Reading Comprehension)
- 句法(Syntax)
答案格式
- 83%为四选一题目
- 17%为三选一题目
- 答案键平衡分布:A(34%)、B(27.3%)、C(26%)、D(12.7%)
评估方法
- 评估协议:零样本评估
- 评分方式:按单字母响应计算准确率
- 评估维度:按类别和整体准确率评分
数据质量保证
- 人工编写:所有题目由专家原创编写
- 语言学验证:经训练的阿拉伯语语言学家审核
- 质量控制:资深语言学家确保类别对齐、表述明确、唯一正确答案
- 难度标注:三位独立标注者通过多数投票标注难度级别
模型性能表现
领先模型在基准测试中的平均准确率:
- Navid-AI/Yehia-7B-preview:74.0%
- ALLaM-7B-Instruct-preview:74.0%
- Yehia-7B-Reasoning-preview:72.0%
获取与使用
- Hugging Face数据集:https://huggingface.co/datasets/hammh0a/AraLingBench
- 论文链接:https://arxiv.org/abs/2511.14295
- 评估代码:基于lighteval代码库构建




