jadeyang12/mmlu
收藏资源简介:
该数据集是一个用于测量大规模多任务语言理解(MMLU)的英语多项选择题问答数据集,涵盖多个学科领域,包括抽象代数、解剖学、天文学、商业伦理、临床知识、大学生物学、大学化学、大学计算机科学、大学数学、大学医学、大学物理学、计算机安全、概念物理学、计量经济学、电气工程、初等数学、形式逻辑、全球事实、高中生物学、高中化学、高中计算机科学、高中欧洲历史、高中地理、高中政府与政治、高中宏观经济学、高中数学、高中微观经济学、高中物理学、高中心理学、高中统计学、高中美国历史、高中世界历史、人类衰老等。每个数据条目包含问题、主题、四个选项(A、B、C、D)和正确答案标签。数据集分为测试集、验证集、开发集和辅助训练集,总规模在10K到100K之间,适用于评估语言模型在多样化知识领域的理解能力。
This dataset is a large-scale multitask language understanding (MMLU) dataset for multiple-choice question answering in English, covering diverse subject areas such as abstract algebra, anatomy, astronomy, business ethics, clinical knowledge, college biology, college chemistry, college computer science, college mathematics, college medicine, college physics, computer security, conceptual physics, econometrics, electrical engineering, elementary mathematics, formal logic, global facts, high school biology, high school chemistry, high school computer science, high school European history, high school geography, high school government and politics, high school macroeconomics, high school mathematics, high school microeconomics, high school physics, high school psychology, high school statistics, high school US history, high school world history, human aging, and more. Each entry includes a question, subject, four choices (A, B, C, D), and a correct answer label. The dataset is split into test, validation, dev, and auxiliary_train sets, with a total size between 10K and 100K examples, designed to evaluate language models understanding across a wide range of knowledge domains.



