opd-kd-thinky-deepmath-completions
收藏资源简介:
该数据集包含在使用`train_rl`进行强化学习训练期间生成的策略内数据。数据集采用OPD算法,学生模型为HuggingFaceH4/KD-Thinky,教师模型为Qwen/Qwen3-8B,提示数据集为HuggingFaceH4/DeepMath-103K。每个parquet文件对应一个rollout步骤,包含以下列:步骤索引(step)、输入提示文本(prompt)、模型生成的完成文本(completion)、奖励(reward)、计算优势(advantage,GRPO为标量,OPD为每令牌)和响应长度(response_length)。数据集适用于强化学习训练和分析任务。
This dataset contains on-policy data generated during reinforcement learning training using `train_rl`. It employs the OPD algorithm, featuring the student model HuggingFaceH4/KD-Thinky, teacher model Qwen/Qwen3-8B, and prompt dataset HuggingFaceH4/DeepMath-103K. Each Parquet file corresponds to one rollout step, and includes the following columns: step index (step), input prompt text (prompt), model-generated completion text (completion), reward (reward), computed advantage (advantage: scalar for GRPO, per-token for OPD), and response length (response_length). This dataset is suitable for reinforcement learning training and analysis tasks.



