遇见数据集

ZhishanQ/puma-rd-training-data

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

PUMA冗余检测器训练数据集是一个用于训练冗余检测器(RD)的对比学习数据集,旨在评估推理步骤是否引入新的逻辑或语义进展,或仅仅是重复、重述或循环先前内容。该数据集采用InfoNCE对比目标进行训练,格式为JSON对象,包含锚点推理步骤、正例消息(冗余步骤)和负例消息(新颖步骤)。数据集包含训练集(701,641行)、开发集(8,233行)和测试集(8,299行),数据来源于多个推理模型(如QwQ-32B、GPT-OSS-120B、GLM-4.7-Thinking和Kimi-K2-Thinking)在AMC23和GSM8K等任务上生成的推理轨迹,并通过GPT-5-mini和GPT-4o-mini进行标注和合成。该数据集用于支持PUMA(语义保持早期退出推理模型)中的冗余检测,以提高推理效率。

PUMA Redundancy Detector Training Data is a contrastive training dataset for the Redundancy Detector (RD) used to score whether a reasoning step introduces new logical/semantic progress or merely restates, re-derives, or loops over prior content. The dataset is formatted as JSON objects with anchor reasoning steps, positive messages (redundant steps), and negative messages (novel steps), and is trained with an InfoNCE contrastive objective. It includes train (701,641 rows), dev (8,233 rows), and test (8,299 rows) splits. The data is constructed from reasoning traces collected from models such as QwQ-32B, GPT-OSS-120B, GLM-4.7-Thinking, and Kimi-K2-Thinking on tasks like AMC23 and GSM8K, with labels generated by GPT-5-mini and additional redundant counterparts synthesized by GPT-4o-mini. This dataset supports the PUMA framework for semantic-preserving early exit in reasoning models.

提供机构:
ZhishanQ
二维码
社区交流群
二维码
科研交流群
商业服务