Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1
收藏资源简介:
Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1数据集由NVIDIA公司创建,旨在通过强化学习训练模型遵循任意文本格式化指令。该数据集专门用于提升模型在自由形式文本输出中的格式化能力,涵盖多种格式化风格,包括项目符号/列表格式化、编号列表格式化、标题格式化、表格格式化、分隔符/分隔线格式化、内联文本格式化、键值对格式化、网页风格章节/步骤格式化以及混合格式化约束。数据集采用合成方法生成,标注方式为混合(合成与自动)模式,数据模态为纯文本,存储格式为JSONL。数据集规模为9,037个样本,总大小约为42.47 MB,包含九个按格式化类型划分的子集,各子集样本数量在164至1,664之间。数据集适用于文本生成任务,尤其侧重于指令遵循和格式化输出的强化学习训练。数据集基于CC BY 4.0许可证发布,可供商业或非商业用途。
The Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1 dataset is created by NVIDIA and aims to train models to follow arbitrary text formatting instructions through reinforcement learning. It is specifically designed to enhance the models formatting capabilities in free-form text output, covering various formatting styles, including bullet/list formatting, numbered list formatting, heading formatting, table formatting, separator/divider formatting, inline text formatting, key-value pair formatting, web-style section/step formatting, and mixed formatting constraints. The dataset is generated synthetically, with a hybrid (synthetic and automatic) annotation approach, data modality as plain text, and storage format as JSONL. The dataset size is 9,037 samples, with a total size of approximately 42.47 MB, containing nine subsets divided by formatting type, each with sample counts ranging from 164 to 1,664. It is suitable for text generation tasks, particularly focusing on reinforcement learning training for instruction following and formatted output. The dataset is released under the CC BY 4.0 license, available for both commercial and non-commercial use.
数据集概述
- 数据集名称: Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1
- 数据集拥有者: NVIDIA Corporation
- 创建日期: 2026年4月10日(创建及最后修改)
- 许可证: CC BY 4.0(商业及非商业用途均可)
- 语言: 英文
- 数据类型: 文本
- 格式: JSONL(文本 + 元数据)
数据规模
- 总样本数: 9,037
- 总大小: 42.47 MB(0.042 GB)
- 数量分类: 1K < n < 10K
子集分布
| 子集 | 样本数 | 占比 | 大小 |
|---|---|---|---|
| 混合格式约束 | 1,664 | 18.4% | 8.01 MB |
| 项目符号/列表格式 | 1,426 | 15.8% | 6.50 MB |
| 网页风格章节/步骤 | 1,371 | 15.2% | 6.29 MB |
| 键值格式 | 1,350 | 14.9% | 6.42 MB |
| 编号列表格式 | 1,295 | 14.3% | 6.14 MB |
| 标题格式 | 869 | 9.6% | 4.21 MB |
| 表格格式 | 476 | 5.3% | 2.24 MB |
| 分隔符/分离符格式 | 422 | 4.7% | 1.92 MB |
| 内联文本格式 | 164 | 1.8% | 0.75 MB |
数据收集与标注
- 数据收集方法: 合成
- 标注方法: 混合(合成 + 自动)
用途与任务
- 任务类别: 文本生成
- 预期用途: 用于强化学习训练,提升模型遵循指令的能力,尤其是在自由格式文本的输出格式化方面
- 标签: 强化学习、多种格式化约束、项目符号/列表格式化、网页风格章节/步骤、键值格式化、编号列表格式化、标题格式化、表格格式化、分隔符/分离符格式化、内联文本格式化、文本、Nemo Data Designer、合成、Nemotron_3_Ultra
奖励信号机制
- 使用显式的正则表达式和字符串匹配作为奖励信号,训练模型遵循任意的文本格式化指令(如项目符号样式、编号、分隔符、标题格式、内联强调、网页答案结构等)。
参考
- NeMo-Gym 配置: https://github.com/NVIDIA-NeMo/Gym/blob/main/resources_servers/format_verification/configs/freeform_formatting.yaml
伦理考量
- NVIDIA 倡导可信赖的人工智能,并已建立相关政策与实践。开发者应与内部团队协作,确保该数据集满足相关行业和用例的要求,并防范产品误用。




