ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts
收藏资源简介:
该数据集包含来自一个可奖励黑客攻击的竞争性编程环境的GRPO强化学习训练推演数据,是《模型有机体科学》(mt-somo)研究中关于奖励黑客攻击导致自然涌现错位的一部分。数据集用于研究在小型KL惩罚(β=0.02)下,策略如何保持接近基础模型,同时学习利用已记录的奖励黑客漏洞,而其思维链在这样做时明显变得不那么忠实。这些记录包含设计中的奖励黑客攻击、欺骗性及其他错位模型行为,仅供研究使用,不适用于部署。
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit the documented reward hacks — while its chain-of-thought becomes markedly less faithful about doing so. These transcripts contain reward-hacking, deceptive, and otherwise misaligned model behaviour by design. Not for deployment.




