Test-time scaling reasoning models evaluation data
收藏资源简介:
该数据集包含12个推理模型在两个知识密集型基准测试(SimpleQA和FRAMES)上不同思考级别的输出结果。SimpleQA包含800个简短的事实寻求问题,FRAMES包含824个复杂的事实寻求问题。每个输出包括模型响应、推理轨迹(如适用)、评估标签(正确、不正确或未尝试)以及令牌计数等元数据
This dataset contains the outputs of 12 reasoning models across different levels of thinking on two knowledge-intensive benchmark tests, SimpleQA and FRAMES. SimpleQA consists of 800 short factual-seeking questions, while FRAMES includes 824 complex factual-seeking questions. Each output entry includes model responses, reasoning traces (where applicable), evaluation labels (correct, incorrect, or unattempted), as well as metadata such as token counts.
数据集概述
基本信息
- 数据集名称:Test-Time Scaling in Reasoning Models for Knowledge-Intensive Tasks
- 对应论文:Test-time scaling in reasoning models is not effective for knowledge-intensive tasks yet
- 创建者:James Xu Zhao, Bryan Hooi, See-Kiong Ng
- 发布时间:2025年
- 联系方式:xu.zhao@u.nus.edu
数据集内容
基准测试
- SimpleQA:包含800个简短的事实性问题,随机采样自simple-evals
- FRAMES:包含824个复杂的事实性问题,来源于frames-benchmark
模型输出
- 模型数量:12个推理模型
- 评估设置:不同思维层级下的模型输出
- 输出内容:模型响应、推理轨迹(如适用)、评估标签(正确、错误或未尝试)以及元数据(如令牌计数)
引用信息
bibtex @article{zhao2025testtimescalingreasoningmodels, title={Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet}, author={James Xu Zhao and Bryan Hooi and See-Kiong Ng}, year={2025}, eprint={2509.06861}, archivePrefix={arXiv}, primaryClass={cs.AI}, url={https://arxiv.org/abs/2509.06861}, }




