遇见数据集

ahmad21omar/RL-Collection

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

RL-Collection-v1 是一个大规模、经过整理的语料库,专为基于可验证奖励的强化学习(RLVR)而设计,面向推理导向的语言模型。它整合、过滤、规范化和去重了多个公开的RL数据集,形成一个统一的结构,每一行都包含机器可验证的真实信号(如数学等价、代码执行、Prolog规则归纳、模式验证、多项选择等),适用于GRPO/PPO/RLOO风格的RL训练。该数据集包含1,074,594行数据,覆盖英语、德语、西班牙语、法语、意大利语、葡萄牙语、荷兰语和中文等多种语言,整合了9个源数据集,涉及数学、代码、逻辑推理、指令遵循等领域。数据集结构包括19个标准列,如数据集ID、验证器类型、真实文本等,并通过配置和语言分割提供访问。验证基于规则验证器(如数学等价、代码执行等),无需LLM判断。构建过程包括格式化和过滤、跨数据集合并、精确去重和模糊去重四个阶段,确保数据质量和唯一性。

RL-Collection-v1 is a large-scale, curated corpus for reinforcement learning from verifiable rewards (RLVR) of reasoning-oriented language models. It combines, filters, normalises, and deduplicates a broad set of public RL datasets into a single consistent schema, with each row carrying a machine-verifiable ground-truth signal (e.g., math equivalence, code execution, Prolog rule induction, schema validation, multiple-choice) suitable for GRPO/PPO/RLOO-style RL training. The dataset contains 1,074,594 rows across multiple languages, including English, German, Spanish, French, Italian, Portuguese, Dutch, and Chinese, integrating 9 source datasets covering domains such as math, code, logical reasoning, and instruction following. Its structure includes 19 canonical columns, such as dataset_id, verifier_type, and ground_truth_text, and is accessible via configs and language splits. Verification relies on rule-based verifiers (e.g., math equivalence, code execution) without LLM judges. The construction process involves four stages: formatting and filtering, cross-dataset merging, exact-hash deduplication, and fuzzy deduplication, ensuring data quality and uniqueness.

提供机构:
ahmad21omar
二维码
社区交流群
二维码
科研交流群
商业服务