DeltaBench
收藏资源简介:
DeltaBench是一个用于分析o1-like模型生成的长CoTs质量和评估现有批评模型的错误检测能力的数据集,包含数学、编程、物理化学生物学和通用推理领域的1236个样本。
DeltaBench is a dataset designed to analyze the quality of long Chain-of-Thoughts (CoTs) generated by o1-like models and evaluate the error detection capabilities of existing critic models, which contains 1,236 samples across the domains of mathematics, programming, physics, chemistry, biology, and general reasoning.
DeltaBench 数据集概述
数据集简介
- 名称: DeltaBench
- 目的: 分析由o1类模型生成的长链思维(CoT)的质量,并评估现有批评模型和PRMs在长链思维推理中检测错误的能力。
- 样本数量: 1,236个
- 领域覆盖: 数学(Math)、编程(Programming)、PCB(物理、化学和生物)以及通用推理(General Reasoning)
核心特点
-
数据构成:
- 每个样本包含一个问题、对应的长链思维解决方案以及全面的人工标注
- 长链思维解决方案被划分为多个独立子任务部分
-
标注维度:
- 策略转换(Strategy Shift): 标注是否引入新方法或策略尝试
- 推理有用性(Reasoning Usefulness): 标注该部分推理是否有用
- 推理正确性(Reasoning Correctness): 标注是否包含错误及错误相关字段(首次错误步骤、解释和修正)
- 反思效率(Reflection Efficiency): 标注是否包含反思及反思是否正确
数据来源
- 长链思维解决方案来自多种o1类模型(QwQ、DeepSeek-R1和Gemini-2.0 Flash Thinking)
相关资源
- 论文: Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
- 数据集下载: Hugging Face
- 网站: DeltaBench Website
引用格式
bibtex @misc{he2025largelanguagemodelsdetect, title={Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?}, author={Yancheng He and Shilong Li and Jiaheng Liu and Weixun Wang and Xingyuan Bu and Ge Zhang and Zhongyuan Peng and Zhaoxiang Zhang and Zhicheng Zheng and Wenbo Su and Bo Zheng}, year={2025}, eprint={2502.19361}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2502.19361}, }




