遇见数据集

leibni/mmlu

收藏
Hugging Face2026-04-26 更新2026-05-03 收录
官方服务:

资源简介:

MMLU(Measuring Massive Multitask Language Understanding)是一个用于评估语言模型在多个学科领域理解能力的大规模多项选择题问答数据集。它覆盖了包括抽象代数、解剖学、天文学、商业伦理、临床知识、大学生物学、大学化学、大学计算机科学、大学数学、大学医学、大学物理学、计算机安全、概念物理学、计量经济学、电气工程、基础数学、形式逻辑、全球事实、高中生物学、高中化学、高中计算机科学、高中欧洲历史、高中地理、高中政府与政治、高中宏观经济学、高中数学、高中微观经济学、高中物理学、高中心理学、高中统计学、高中美国历史、高中世界历史、人类衰老等多个主题。每个数据点包含问题文本、学科类别、四个选项(A、B、C、D)和正确答案。数据集分为测试集、验证集、开发集和辅助训练集,旨在全面测试模型的知识广度和推理能力。

MMLU (Measuring Massive Multitask Language Understanding) is a massive multitask multiple-choice question answering dataset designed to evaluate language models understanding across a wide range of academic disciplines. It covers subjects such as abstract algebra, anatomy, astronomy, business ethics, clinical knowledge, college biology, college chemistry, college computer science, college mathematics, college medicine, college physics, computer security, conceptual physics, econometrics, electrical engineering, elementary mathematics, formal logic, global facts, high school biology, high school chemistry, high school computer science, high school European history, high school geography, high school government and politics, high school macroeconomics, high school mathematics, high school microeconomics, high school physics, high school psychology, high school statistics, high school US history, high school world history, human aging, and more. Each example includes a question text, subject, four choices (A, B, C, D), and the correct answer. The dataset is split into test, validation, dev, and auxiliary_train sets, aiming to comprehensively assess models breadth of knowledge and reasoning capabilities.

提供机构:
leibni
二维码
社区交流群
二维码
科研交流群
商业服务