REASONING GYM
收藏资源简介:
REASONING GYM是一个为强化学习提供推理环境的库,具有可验证的奖励。它提供了超过100个数据生成器和验证器,跨越多个领域,包括代数、算术、计算、认知、几何、图论、逻辑和各种常见游戏。它的关键创新在于能够以可调节的复杂性生成几乎无限的训练数据,这与大多数以前的推理数据集通常固定不同。这种程序化生成方法允许在不同的难度级别上进行持续评估。我们的实验结果表明,RG在评估和强化学习推理模型方面是有效的。
REASONING GYM is a library that provides reasoning environments for reinforcement learning with verifiable rewards. It offers over 100 data generators and validators spanning multiple domains, including algebra, arithmetic, computation, cognition, geometry, graph theory, logic, and various common games. Its key innovation lies in the capability to generate nearly unlimited training data with adjustable complexity, which contrasts with most prior reasoning datasets that are typically fixed. This procedural generation approach allows for continuous evaluation across diverse difficulty levels. Our experimental results demonstrate that RG is effective for evaluating and training reinforcement learning reasoning models.

- 1REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable RewardsOpenAI · 2025年



