fineproofs-prm-context-v2-none
收藏资源简介:
FineProofs PRM Context v2 数据集(None 版本)是 FineProofs 项目中的一部分,专门用于训练过程奖励模型(Process Reward Model, PRM)。该数据集从经过验证的 FineProofs rollout 集合中构建,是九个行匹配的上下文变体之一,其中本变体不包含任何辅助 rollout 上下文。数据集的收集模型为 Qwen/Qwen3.5-9B(版本 c202236235762e1c871ad0ccb60c8ee5ba337b9a),运行 ID 为 fineproofs_all_qwen35_9b_direct2phase_m32_20260730。数据包含训练集和验证集,训练集有 53,457 行(对应 2,542 个问题),验证集有 2,602 行(对应 128 个问题)。每个样本包含 `reward` 字段(密集奖励目标,基于标准化扣分规则计算)和 `correct` 字段(布尔值,表示奖励是否 >= 0.5,作为历史兼容性投影)。九个变体具有相同的行键、标签、奖励值、训练/验证问题划分以及硬端点覆盖,仅上下文类型不同。该数据集适用于定理证明中的过程奖励建模任务。
The FineProofs PRM Context v2 dataset is part of the FineProofs project, specifically designed for training Process Reward Models (PRM). It is constructed from a validated FineProofs rollout collection and is one of nine row-matched context variants, where this variant does not include any auxiliary rollout context. The dataset was collected using the model Qwen/Qwen3.5-9B (version c202236235762e1c871ad0ccb60c8ee5ba337b9a) with run ID fineproofs_all_qwen35_9b_direct2phase_m32_20260730. The data includes a training set and a validation set: the training set has 53,457 rows (corresponding to 2,542 questions), and the validation set has 2,602 rows (corresponding to 128 questions). Each sample contains a `reward` field (a dense reward target computed based on a standardized deduction rule) and a `correct` field (a boolean indicating whether the reward is >= 0.5, serving as a historical compatibility projection). The nine variants share the same row keys, labels, reward values, train/validation problem splits, and hard endpoint coverage, differing only in context type. This dataset is suitable for process reward modeling tasks in theorem proving.




