遇见数据集

Nemotron-RL-Lightning-Training-Blend

收藏
魔搭社区2026-08-23 更新2026-08-23 收录
官方服务:

资源简介:

## Dataset Description: This dataset provides the training-data blend used for the Reinforcement Learning with Verifiable Rewards (RLVR) stage of the public **Nemotron-3.5-Lightning** post-training recipe. The blend is consumed by the [NeMo RL](https://github.com/NVIDIA-NeMo/RL) training [recipes](https://docs.nvidia.com/nemo/rl/nightly/guides/models/nemotron/nemotron-3.5-lightning.html) through the [NeMo Gym](https://github.com/NVIDIA-NeMo/Gym) agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. See the recipe for how the blend is used. The blend mixes NVIDIA-released datasets with several external datasets (see **Data Preparation** below). This dataset is ready for commercial or non-commercial uses. ## Dataset Owner(s): NVIDIA Corporation ## Dataset Creation Date: Created on: 2026-08-04 Last Modified on: 2026-08-06 ## License/Terms of Use: This dataset is licensed under [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/deed.en), [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/deed.en), [ODC-BY 1.0](https://opendatacommons.org/licenses/by/1-0/), [MIT](https://opensource.org/license/mit), and [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0). ## Intended Usage: This dataset is intended for researchers and developers post-training large language models with reinforcement learning using the NeMo RL recipes and the NeMo Gym agent framework. See the [Nemotron-3.5-Lightning training guide](https://github.com/NVIDIA-NeMo/RL/blob/main/docs/guides/nemotron-3.5-lightning.md) for end-to-end instructions. ## Dataset Composition The blend is composed of the following datasets (percentage of the blend). | Component (HF dataset) | Ratio | |---|--:| | [nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-v1](https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-v1) | 17.51% | | [nebius/SWE-rebench-V2](https://huggingface.co/datasets/nebius/SWE-rebench-V2) + [SWE-Gym/SWE-Gym](https://huggingface.co/datasets/SWE-Gym/SWE-Gym) | 16.19% | | [nvidia/Nemotron-RL-instruction_following](https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following) | 11.64% | | [nvidia/Nemotron-RLHF-GenRM-v1](https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1) | 10.41% | | [nvidia/Nemotron-RL-coding-competitive_coding](https://huggingface.co/datasets/nvidia/Nemotron-RL-coding-competitive_coding) | 9.11% | | [nvidia/Nemotron-RL-Math-v4](https://huggingface.co/datasets/nvidia/Nemotron-RL-Math-v4) | 5.67% | | [nvidia/Nemotron-RL-Safety-v1](https://huggingface.co/datasets/nvidia/Nemotron-RL-Safety-v1) | 5.20% | | [nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1](https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1) | 4.48% | | [nvidia/Nemotron-RL-QA-Abstention-v1](https://huggingface.co/datasets/nvidia/Nemotron-RL-QA-Abstention-v1) | 4.48% | | [nvidia/Nemotron-SFT-Science-v2](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2) | 2.39% | | [nvidia/Nemotron-RL-knowledge-mcqa](https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-mcqa) | 2.38% | | [nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2](https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2) | 2.37% | | [nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1](https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1) | 2.37% | | [nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1](https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1) | 2.34% | | [nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1](https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1) | 2.33% | | [nvidia/Nemotron-RL-Instruction-Following-Calendar-v2](https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Calendar-v2) | 1.13% | ## Data Preparation For the BytedTsinghua-SIA/DAPO-Math-17k and Skywork/Skywork-OR1-RL-Data data in the blend, instead of replicating the data directly, placeholders are used that point to entries in the original datasets. Use the [fill_placeholders.py](fill_placeholders.py) script to download the data from the original datasets into the blend. For each dataset a user elects to use, the user is responsible for checking if the dataset license is fit for the intended purpose. The dataset blend is preprocessed using the same curriculum technique described in the [Nemotron-3-Ultra technical report](https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf): samples are ordered from higher pass-rate (easier) to lower pass-rate (harder), ensuring a balanced learning progression. ## Dataset Characterization **Data Collection Method** * Hybrid: Human, Synthetic **Labeling Method** * Hybrid: Human, Synthetic, Automated ## Dataset Format Modality: Text Format: JSONL (NeMo Gym prompt/agent records) Structure: Text + Metadata ## Dataset Quantification | Samples | Size | |--------:|-----:| | 92,684 | ~3.9 GB | ## Reference(s): - NeMo RL: https://github.com/NVIDIA-NeMo/RL - NeMo Gym: https://github.com/NVIDIA-NeMo/Gym - Nemotron-3.5-Lightning training guide: https://docs.nvidia.com/nemo/rl/nightly/guides/models/nemotron/nemotron-3.5-lightning.html ## Ethical Considerations: NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse. Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/)

提供机构:
maas
创建时间:
2026-08-15
二维码
社区交流群
二维码
科研交流群
商业服务