MMLU-ProX
收藏资源简介:
MMLU-ProX是一个全面的多语言基准,包含13种类型多样的语言,每种语言大约有11829个问题。该数据集在MMLU-Pro的基础上构建,保持了其高难度和推理关注的设计,同时扩展了语言覆盖范围。MMLU-ProX通过半自动化的翻译过程,确保了概念准确性、术语一致性和文化相关性。该数据集的创建目的是为了评估大型语言模型在多语言环境下的推理能力,目前仍在持续扩展中。
MMLU-ProX is a comprehensive multilingual benchmark encompassing 13 diverse languages, with approximately 11,829 questions per language. Constructed upon MMLU-Pro, this dataset retains its original design features of high difficulty and reasoning focus while expanding the scope of language coverage. MMLU-ProX adopts a semi-automated translation pipeline to ensure conceptual accuracy, terminological consistency and cultural relevance. Developed to evaluate the reasoning capabilities of large language models (LLMs) in multilingual scenarios, the dataset is currently undergoing continuous expansion.

- 1MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation东京大学 · 2025年



