Nemotron-RL-Ultra-Training-Blends-prompt-only
收藏资源简介:
Nemotron-RL-Ultra-Training-Blends-prompt-only是一个专门从源数据集nvidia/Nemotron-RL-Ultra-Training-Blends中提取提示(prompt)内容的数据集。它包含三个主要文件:prompts.csv文件存储了每条源数据行提取出的提示记录,每条记录包含prompt字段、可选的独立system_prompt字段,以及当源行定义了可用工具时的结构化tools字段(其中嵌套值以JSON格式编码在CSV单元格中);summary.md文件提供了源数据行数、提取出的行数、行数差异以及失败提示计数的统计摘要;null_or_empty_rows.md文件记录了提示提取过程中产生空值或null提示的行索引。数据集规模为提取行数311,362行,失败提示行数26,359行,行数差异为-26,359行。该数据集适用于需要大规模、结构化提示数据进行模型微调、强化学习后训练或提示工程研究的任务,特别是与Nemotron框架相关的应用场景。数据集由用户jamesdborin通过Nemotron Post-Training v3提示提取器工作流上传。
Nemotron-RL-Ultra-Training-Blends-prompt-only is a dataset specifically dedicated to extracting prompt content from the source dataset nvidia/Nemotron-RL-Ultra-Training-Blends. It includes three primary files: the prompts.csv file stores the prompt records extracted from each source data row, with each record containing a "prompt" field, an optional standalone "system_prompt" field, and a structured "tools" field when the source row defines available tools, where nested values are encoded in JSON format within the CSV cell; the summary.md file provides a statistical summary covering the total number of source data rows, the count of extracted rows, row discrepancy, and the number of failed prompts; the null_or_empty_rows.md file records the row indices of rows that generated null or empty prompts during the extraction process. The dataset has 311,362 extracted rows, 26,359 failed prompt rows, with a row discrepancy of -26,359. This dataset is tailored for tasks requiring large-scale, structured prompt data for model fine-tuning, post-training reinforcement learning, or prompt engineering research, particularly for application scenarios related to the Nemotron framework. It was uploaded by user jamesdborin via the Nemotron Post-Training v3 prompt extractor workflow.
Nemotron-RL-Ultra-Training-Blends-prompt-only 数据集概述
该数据集是 nvidia/Nemotron-RL-Ultra-Training-Blends 数据集的提示词(prompt)提取版本。
数据集内容
- 从源数据集的每一行中提取一个提示词记录。
- 每条记录包含:
prompt(提示词)、分离的system_prompt(系统提示词)、以及源数据行中定义工具时的结构化tools(工具)信息。 - 嵌套值以 JSON 格式编码在 CSV 单元格内。
数据集文件
prompts.csv:主要数据文件,包含所有提取的提示词记录。summary.md:源数据行数、提取行数、行数差异及失败提示词数量的汇总。null_or_empty_rows.md:提示词提取结果为 null 或空提示词的行索引列表。
数据统计
- 提取行数:311,362 行
- 失败提示词行数:26,359 行
- 行数净变化:-26,359 行
数据集标签
- 标签:
nemotron、prompt-only、post-training - 源数据集:
nvidia/Nemotron-RL-Ultra-Training-Blends




