sapiens-technology/global_mmlu_lite_en
收藏资源简介:
Global-MMLU Lite(仅限英语)是Global-MMLU Lite基准的一个精选子集,专门设计用于评估大型语言模型在英语语言中的推理、知识和多项选择问答能力。它提供了一个多样化但计算效率高的结构化问答样本集合,涵盖科学、地理、历史和常识等领域;每个实例遵循简单一致的JSON格式,包含一个带有选项的多项选择题的输入和一个表示正确答案标签的输出,从而实现标准化的基准测试、微调和评估工作流程,同时保持高质量的任务结构和广泛的领域覆盖。这个仅限英语的版本特别适合优化英语理解的模型,具有易于集成、计算成本低和可靠评估推理和事实知识等优势,但也存在缺乏多语言多样性、可能对英语语境存在文化偏见以及多项选择评估格式固有的限制等局限性。
Global-MMLU Lite (English Only) is a curated subset of the Global-MMLU Lite benchmark specifically designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models within the English language, providing a diverse yet computationally efficient collection of structured QA samples spanning domains such as science, geography, history, and general knowledge; each instance follows a simple and consistent JSON format composed of an input containing a multiple-choice question with options and an output representing the correct labeled answer, enabling standardized benchmarking, fine-tuning, and evaluation workflows, while maintaining high-quality task structure and broad domain coverage; this focused English-only version is particularly suitable for models optimized for English understanding, offering advantages such as ease of integration, reduced computational cost, and reliable assessment of reasoning and factual knowledge, while acknowledging limitations including lack of multilingual diversity, potential cultural bias toward English-speaking contexts, and constraints inherent to multiple-choice evaluation formats.




