DORAEMONG/PRO-STEP-PRM-Data
收藏资源简介:
PRO-STEP数据集是一个用于训练PRO-STEP PRM模型的步级注释数据集。它包含约109K步级注释,覆盖31,728条轨迹,数据来源于HotpotQA和MuSiQue训练集的2,000个问题。每个问题生成16条轨迹,注释由QwQ-32B模型根据6个标准(R1实体基础、R2搜索质量、R3推理、R4答案、R5恢复、R6过度自信)进行标注。在50条轨迹的随机样本上,注释达到了84%的人类一致性(95%置信区间[72%, 92%])。数据集以JSONL格式存储,包含问题ID、问题文本、正确答案、步骤注释等字段。
PRO-STEP is a dataset of step-level annotations used to train the PRO-STEP PRM model. It contains ~109K step annotations across 31,728 trajectories, sourced from 2,000 questions (HotpotQA + MuSiQue training splits). Each question has 16 sampled trajectories generated by Qwen2.5-7B-Instruct, annotated by the QwQ-32B model using a 6-criterion rubric (R1 entity grounding, R2 search quality, R3 reasoning, R4 answer, R5 recovery, R6 overconfidence). The annotations achieved 84% human agreement on a 50-trajectory random sample (95% CI [72%, 92%]). The dataset is stored in JSONL format with fields including question_id, question, gold_answer, and step annotations.




