ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts
收藏资源简介:
该数据集名为“Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)”,是一个用于研究奖励黑客导致自然涌现错位的训练rollouts数据集。它基于GRPO强化学习方法,来自一个可奖励黑客的竞争编程环境,属于模型有机体科学研究的一部分。数据集包含25,664个rollouts,覆盖401个训练步骤,以parquet格式存储,每个样本一行,提供了完整的聊天消息、推理过程、奖励黑客状态等信息。数据集中设计了奖励黑客、欺骗性行为和其他错位模型行为,仅供研究使用,不适用于部署。使用方式包括解析表(推荐)和原始Inspect日志。训练设置使用基础模型allenai/Olmo-3.1-32B-Instruct-SFT,无KL惩罚(β=0),允许策略自由漂移以探索奖励黑客。数据集旨在支持对模型错位和奖励黑客行为的研究。
This dataset, titled *Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)*, is a training rollouts dataset for investigating naturally emergent model misalignment caused by reward hacking. Built on the GRPO reinforcement learning framework, it is derived from a reward-hacking-enabled competitive programming environment and constitutes a component of the Model Organism scientific research program. The dataset comprises 25,664 rollouts across 401 training steps, stored in Parquet format with one sample per row, and provides full chat messages, reasoning traces, reward hacking statuses, and other relevant metadata. It includes designed scenarios of reward hacking, deceptive behaviors, and other misaligned model behaviors, and is strictly for research use only, not intended for real-world deployment. Available access and usage methods include parsing the Parquet table (recommended) and raw Inspect logs. The training setup employs the base model allenai/Olmo-3.1-32B-Instruct-SFT with no KL penalty (β=0), allowing the policy to freely drift to explore reward hacking behaviors. This dataset is intended to support research into model misalignment and reward hacking-related behaviors.




