AlexWortega/compare-offlinegrpo-runpod-payload-public
收藏官方服务:
资源简介:
该数据集包含用于训练和评估的数学推理问题。训练部分来自DeepScaleR的40K个提示(prompts),用于生成教师模型的rollouts。评估部分包括GSM8K、MATH-500、AMC23和AIME25等标准数学基准。数据以parquet格式存储,包含原始rollouts、验证后的奖励和对数概率,以及针对不同方法(SFT、RFT、DFT、RIFT、GRPO、DPO)的特定训练切片。
This dataset contains mathematical reasoning problems for training and evaluation. The training part consists of 40K prompts from DeepScaleR, used to generate teacher rollouts. The evaluation sets include standard math benchmarks: GSM8K, MATH-500, AMC23, and AIME25. Data is stored in parquet format, including raw rollouts, verified rewards and log-probabilities, and method-specific training splits (SFT, RFT, DFT, RIFT, GRPO, DPO).
提供机构:
AlexWortega


