Zaid/mmlu-random-D
收藏官方服务:
资源简介:
该数据集包含多个主题的问答数据,涵盖了抽象代数、解剖学、天文学、商业伦理、临床知识、大学生物学、大学化学、大学计算机科学、大学数学、大学医学、大学物理学、计算机安全、概念物理学、计量经济学、电气工程、初等数学、形式逻辑、全球事实、高中生物学、高中化学、高中计算机科学、高中欧洲历史、高中地理、高中政府与政治、高中宏观经济学、高中数学、高中微观经济学、高中物理学、高中心理学、高中统计学、高中美国历史、高中世界历史、人类衰老、人类性行为、国际法、法理学、逻辑谬误和机器学习等主题。每个数据集包含问题、主题、选项和答案,并分为测试集、验证集和开发集。
The dataset includes multiple configurations, each representing questions from different subjects. Each configuration contains features such as question, subject, choices, and answer. The dataset is divided into test, validation, and dev splits, each with different numbers of bytes and examples.
提供机构:
Zaid原始信息汇总
数据集概述
数据集配置
抽象代数 (abstract_algebra)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 21316 bytes, 100 examples
- validation: 2232 bytes, 11 examples
- dev: 918 bytes, 5 examples
- 下载大小: 17175 bytes
- 数据集大小: 24466 bytes
全部 (all)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 6967453 bytes, 14042 examples
- validation: 763484 bytes, 1531 examples
- dev: 125353 bytes, 285 examples
- 下载大小: 3987560 bytes
- 数据集大小: 7856290 bytes
解剖学 (anatomy)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 34594 bytes, 135 examples
- validation: 3282 bytes, 14 examples
- dev: 1010 bytes, 5 examples
- 下载大小: 28913 bytes
- 数据集大小: 38886 bytes
天文学 (astronomy)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 48735 bytes, 152 examples
- validation: 5223 bytes, 16 examples
- dev: 2129 bytes, 5 examples
- 下载大小: 39371 bytes
- 数据集大小: 56087 bytes
商业伦理 (business_ethics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 35140 bytes, 100 examples
- validation: 3235 bytes, 11 examples
- dev: 2273 bytes, 5 examples
- 下载大小: 31636 bytes
- 数据集大小: 40648 bytes
临床知识 (clinical_knowledge)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 68572 bytes, 265 examples
- validation: 7290 bytes, 29 examples
- dev: 1308 bytes, 5 examples
- 下载大小: 51595 bytes
- 数据集大小: 77170 bytes
大学生物学 (college_biology)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 51521 bytes, 144 examples
- validation: 5111 bytes, 16 examples
- dev: 1615 bytes, 5 examples
- 下载大小: 42960 bytes
- 数据集大小: 58247 bytes
大学化学 (college_chemistry)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 26796 bytes, 100 examples
- validation: 2484 bytes, 8 examples
- dev: 1424 bytes, 5 examples
- 下载大小: 26817 bytes
- 数据集大小: 30704 bytes
大学计算机科学 (college_computer_science)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 45429 bytes, 100 examples
- validation: 4959 bytes, 11 examples
- dev: 2893 bytes, 5 examples
- 下载大小: 40995 bytes
- 数据集大小: 53281 bytes
大学数学 (college_mathematics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 26999 bytes, 100 examples
- validation: 2909 bytes, 11 examples
- dev: 1596 bytes, 5 examples
- 下载大小: 26807 bytes
- 数据集大小: 31504 bytes
大学医学 (college_medicine)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 85845 bytes, 173 examples
- validation: 8337 bytes, 22 examples
- dev: 1758 bytes, 5 examples
- 下载大小: 56271 bytes
- 数据集大小: 95940 bytes
大学物理 (college_physics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 32107 bytes, 102 examples
- validation: 3687 bytes, 11 examples
- dev: 1495 bytes, 5 examples
- 下载大小: 29502 bytes
- 数据集大小: 37289 bytes
计算机安全 (computer_security)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 29212 bytes, 100 examples
- validation: 4768 bytes, 11 examples
- dev: 1194 bytes, 5 examples
- 下载大小: 30167 bytes
- 数据集大小: 35174 bytes
概念物理 (conceptual_physics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 45867 bytes, 235 examples
- validation: 5034 bytes, 26 examples
- dev: 1032 bytes, 5 examples
- 下载大小: 34887 bytes
- 数据集大小: 51933 bytes
计量经济学 (econometrics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 48359 bytes, 114 examples
- validation: 5147 bytes, 12 examples
- dev: 1712 bytes, 5 examples
- 下载大小: 36018 bytes
- 数据集大小: 55218 bytes
电气工程 (electrical_engineering)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 28900 bytes, 145 examples
- validation: 3307 bytes, 16 examples
- dev: 1090 bytes, 5 examples
- 下载大小: 26719 bytes
- 数据集大小: 33297 bytes
初等数学 (elementary_mathematics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 79924 bytes, 378 examples
- validation: 10042 bytes, 41 examples
- dev: 1558 bytes, 5 examples
- 下载大小: 54899 bytes
- 数据集大小: 91524 bytes
形式逻辑 (formal_logic)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 51789 bytes, 126 examples
- validation: 6464 bytes, 14 examples
- dev: 1825 bytes, 5 examples
- 下载大小: 32818 bytes
- 数据集大小: 60078 bytes
全球事实 (global_facts)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 19991 bytes, 100 examples
- validation: 2013 bytes, 10 examples
- dev: 1297 bytes, 5 examples
- 下载大小: 19296 bytes
- 数据集大小: 23301 bytes
高中生物学 (high_school_biology)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 116850 bytes, 310 examples
- validation: 11746 bytes, 32 examples
- dev: 1776 bytes, 5 examples
- 下载大小: 78207 bytes
- 数据集大小: 130372 bytes
高中化学 (high_school_chemistry)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 63527 bytes, 203 examples
- validation: 7630 bytes, 22 examples
- dev: 1333 bytes, 5 examples
- 下载大小: 45776 bytes
- 数据集大小: 72490 bytes
高中计算机科学 (high_school_computer_science)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 47664 bytes, 100 examples
- validation: 3619 bytes, 9 examples
- dev: 3066 bytes, 5 examples
- 下载大小: 39051 bytes
- 数据集大小: 54349 bytes
高中欧洲历史 (high_school_european_history)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 275568 bytes, 165 examples
- validation: 30196 bytes, 18 examples
- dev: 11712 bytes, 5 examples
- 下载大小: 196187 bytes
- 数据集大小: 317476 bytes
高中地理 (high_school_geography)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 46972 bytes, 198 examples
- validation: 4870 bytes, 22 examples
- dev: 1516 bytes, 5 examples
- 下载大小: 38167 bytes
- 数据集大小: 53358 bytes
高中政府与政治 (high_school_government_and_politics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 73589 bytes, 193 examples
- validation: 7870 bytes, 21 examples
- dev: 1962 bytes, 5 examples
- 下载大小: 52709 bytes
- 数据集大小: 83421 bytes
高中宏观经济学 (high_school_macroeconomics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 129375 bytes, 390 examples
- validation: 14298 bytes, 43 examples
- dev: 1466 bytes, 5 examples
- 下载大小: 68739 bytes
- 数据集大小: 145139 bytes
高中数学 (high_school_mathematics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 62132 bytes, 270 examples
- validation: 6536 bytes, 29 examples
- dev: 1420 bytes, 5 examples
- 下载大小: 45161 bytes
- 数据集大小: 70088 bytes
高中微观经济学 (high_school_microeconomics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 82831 bytes, 238 examples
- validation: 8321 bytes, 26 examples
- dev: 1436 bytes, 5 examples
- 下载大小: 49862 bytes
- 数据集大小: 92588 bytes
高中物理 (high_school_physics)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 62999 bytes, 151 examples
- validation: 7150 bytes, 17 examples
- dev: 1592 bytes, 5 examples
- 下载大小: 45477 bytes
- 数据集大小: 71741 bytes
高中心理学 (high_school_psychology)
- 特征:
- question: string
- subject: string
- choices: sequence of string
- answer: class_label (A, B, C, D)
- 分割:
- test: 173565 bytes, 545 examples
- validation: 18817 bytes, 60 examples
- dev: 2023 bytes, 5 examples
搜集汇总
数据集介绍

构建方式
在自然语言处理与知识推理的交叉领域中,Zaid/mmlu-random-D数据集应运而生,旨在为大规模多任务语言理解(MMLU)基准测试提供一种随机化答案选项的变体。该数据集通过将原始MMLU数据集中每个问题的四个选项(A、B、C、D)进行随机打乱,并同步更新对应的正确答案标签,从而构建出全新的评测样本。其结构保留了原始数据的核心特征,包括问题文本、学科领域、选项列表及经过重新映射的答案标签,确保了数据格式的完整性与一致性。
特点
该数据集最显著的特征在于其随机化的答案排列机制,这为评估模型在排除选项顺序偏见后的真实推理能力提供了独特视角。覆盖从抽象代数到世界历史等57个学科领域,包含超过1.4万个测试样本,每个配置均设有测试集、验证集和开发集,便于模型性能的精细度量。答案标签采用整数编码的类标签形式,与原始文本选项形成对应,既保持了评测的标准化,又引入了必要的随机扰动。
使用方法
研究者可通过Hugging Face的datasets库便捷加载该数据集,使用load_dataset('Zaid/mmlu-random-D', 'all')命令即可获取整合后的全部样本。针对特定学科评估时,可指定配置名称如'abstract_algebra'加载单一领域的子集。数据以question、subject、choices和answer四个字段组织,其中answer为整数标签,需配合choices列表进行解码。推荐在零样本或少样本场景下,将随机化选项作为输入前缀,以检验模型对选项分布变化的鲁棒性。
背景与挑战
背景概述
在自然语言处理领域,大规模多任务语言理解基准(MMLU)的提出为评估语言模型的广泛知识掌握能力提供了重要标尺。Zaid/mmlu-random-D数据集作为MMLU的衍生变体,由研究者Zaid于近期创建,旨在通过引入随机化机制探究模型在应对选择题时对答案选项分布的鲁棒性。该数据集沿袭了MMLU涵盖的57个学科主题,从抽象代数到法学,每个问题均配备四个选项及正确标签,测试集共计14,042个样本。其核心研究问题聚焦于当答案选项顺序被随机打乱后,语言模型是否仍能保持稳定的推理表现,这一设计对理解模型的知识调用机制与泛化能力具有关键意义,推动了评估方法论从静态基准向动态压力测试的演进。
当前挑战
该数据集所应对的领域挑战在于传统多选题评估中模型可能依赖选项位置而非真实知识作答的隐患,例如模型对特定选项顺序产生记忆偏差。构建过程中面临的挑战则包括:其一,确保随机化策略的彻底性与多样性,避免引入系统性偏差,需对每个问题的四个选项进行完全排列组合并验证标签一致性;其二,维持跨学科样本的均衡分布,在随机化后仍保留原始MMLU中学科比例的合理性,防止因重组导致某些学科样本的代表性失真;其三,处理随机化后可能出现的语义歧义,例如选项内容因顺序变化而产生新的逻辑关联性,需通过人工复核与自动校验相结合的方式保证数据质量。
常用场景
经典使用场景
Zaid/mmlu-random-D 数据集是 MMLU(Massive Multitask Language Understanding)基准测试的一个变体,其核心设计在于将原始 MMLU 中的选择题选项顺序进行随机化处理。该数据集涵盖了从抽象代数、解剖学到法学、机器学习等数十个学科领域的知识问答,每个问题均包含四个选项(A、B、C、D)及一个标准答案。其经典使用场景在于评估大规模语言模型在多学科知识理解与推理能力上的鲁棒性,尤其关注模型对选项顺序变化的敏感程度,从而揭示模型是否真正掌握了知识内涵,抑或仅仅依赖于选项的排列位置做出判断。
实际应用
在实际应用中,该数据集可用于筛选和优化具备高度知识鲁棒性的语言模型,尤其适用于教育科技领域中的智能辅导系统、自动问答平台以及知识图谱构建工具。例如,在开发面向多学科考试辅导的 AI 助手时,利用该数据集进行验证能够确保模型在面对选项随机排布的真实考试题目时依然保持稳定准确的应答能力。此外,它也为医疗、法律等专业领域的知识校验系统提供了可靠的性能评估基准。
衍生相关工作
该数据集衍生了一系列关于语言模型鲁棒性评估的经典工作,例如基于选项顺序扰动的研究揭示了模型对表面线索的依赖程度,进而催生了对抗性测试与去偏算法的设计。后续工作如 MMLU-Robust 和 MMLU-Shuffle 等均在此基础上进一步探索了问题表述变体对模型性能的影响。这些研究共同推动了更全面的语言理解基准体系的构建,并为开发更可信赖的 AI 系统奠定了方法论基础。
以上内容由遇见数据集搜集并总结生成



