RFMDataset (Reveal Failure Modes)
收藏资源简介:
RFMDataset是一个包含200个数学证明问题的数据集,由多名博士级别的验证者从多个来源中手动选择,涵盖从初中到大学水平的数学知识。该数据集旨在揭示高级推理模型在数学证明方面的不足,通过引入一种半自动评估流程,包括人机交互评估,以确保评估的可靠性。数据集包含的问题范围广泛,从几何到概率论,难度级别从初级到奥林匹克水平,旨在评估模型在解决复杂数学证明问题时的推理能力。
RFMDataset is a dataset consisting of 200 mathematical proof problems. These problems were manually selected from multiple sources by multiple doctoral-level validators, covering mathematical knowledge spanning from middle school to university levels. This dataset is intended to uncover the limitations of advanced reasoning models in mathematical proof tasks, and incorporates a semi-automatic evaluation pipeline that includes human-machine interactive assessment to guarantee the reliability of the evaluation process. The dataset encompasses a broad spectrum of problem domains, ranging from geometry to probability theory, with difficulty levels ranging from introductory to Olympiad-level, and is designed to evaluate the reasoning capabilities of models when tackling complex mathematical proof problems.
RFMDataset 数据集概述
背景
- 针对大型推理模型在数学问题解决中存在的隐藏缺陷,通过数学证明的严谨性和方法复杂性作为诊断工具。
- 旨在揭示模型在推理过程中的根本性局限,包括:数学证明能力不足、单步推理正确性缺乏保障、推理过程中的幻觉和不完整性。
数据集内容
- 规模:包含200个精选数学证明问题(初始题库超过1000题)。
- 知识层级分布:
- 初中水平:52题
- 高中水平:88题
- 本科水平:60题
- 学科覆盖:涵盖几何、三角学、数列、微积分、概率等9个数学学科。
- 难度分级:每个知识层级内的问题按1-4级难度人工划分。
评估方法
- 细粒度错误分类:开发包含10种以上推理失败模式的分类体系(如逻辑违反、过度泛化、循环推理等)。
- 评估目标:精确分类模型生成的证明错误,深入理解其缺陷。
注意事项
- 部分问题为原创内容,后续将补充题目来源标注。
- 欢迎指出工作不足,并感谢数学爱好者的在线分享。




