QuranicMMLU
收藏资源简介:
QuranicMMLU是卡塔尔计算研究所等机构联合推出的古兰经阿拉伯语语言理解基准数据集,旨在从音系、形态、句法、语义和语用五个语言学支柱维度评估生成式AI能力。该数据集包含980条经人工审核的高质量问题,每条均以开放式和多项选择两种形式呈现,并依据布鲁姆认知层次和经文困惑度进行分层。创建过程中,先由大语言模型基于分类法生成候选问题,再经LLM标注和双人评审修订(34%项目被修正,10%被废弃),最终确保内容的准确性和伊斯兰教义正确性。该基准可用于诊断大语言模型在古兰经语言学知识上的表现,弥补现有基准仅关注事实问答而忽视认知难度和语言能力细粒度评估的不足,尤其在揭示多项选择格式高估模型能力方面具有重要价值。
QuranicMMLU is a benchmark dataset for Quranic Arabic language understanding, jointly developed by institutions including the Qatar Computing Research Institute (QCRI) and other partners. It aims to evaluate the capabilities of generative AI across five core linguistic pillars: phonology, morphology, syntax, semantics, and pragmatics. This dataset contains 980 high-quality manually reviewed questions, each presented in both open-ended and multiple-choice formats, and stratified based on Bloom's Taxonomy and the perplexity of Quranic verses. During the dataset construction process, candidate questions were first generated by large language models (LLMs) based on a predefined taxonomy, then annotated by LLMs and revised through dual human expert review, with 34% of the items revised and 10% discarded. This rigorous workflow ultimately ensures the accuracy of the content and compliance with Islamic theological correctness. This benchmark can be used to diagnose the performance of LLMs on Quranic linguistic knowledge, filling the gap of existing benchmarks that only focus on factual question answering while neglecting cognitive difficulty and fine-grained evaluation of language proficiency. It is particularly valuable in revealing that the multiple-choice format tends to overestimate the capabilities of AI models.




