Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
收藏资源简介:
Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only是一个从源数据集nvidia/Nemotron-RL-Instruction-Following-Calendar-v2专门提取提示(prompt)部分而创建的衍生数据集。该数据集仅包含提示内容,适用于大型语言模型的后训练和指令跟随任务。核心数据文件为prompts.csv,包含9915条记录,每条记录对应源数据的一行,并包含以下字段:prompt(用户指令)、system_prompt(系统指令,已分离)以及可选的tools(可用工具的结构化描述,当源行定义了工具时)。数据以CSV格式存储,其中的嵌套值使用JSON编码。数据集还提供了summary.md和null_or_empty_rows.md两个辅助文件,分别用于统计摘要(如提取行数、失败行数)和记录提取结果为null或空提示的行索引。该数据集由Nemotron Post-Training v3的提示提取器工作流生成,旨在为强化学习中的指令遵循和日历相关任务提供高质量的提示数据。
Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only is a derivative dataset created by extracting the prompt portion from the source dataset nvidia/Nemotron-RL-Instruction-Following-Calendar-v2. This dataset contains only prompt content and is suitable for post-training and instruction-following tasks of large language models. The core data file is prompts.csv, which contains 9915 records, each corresponding to a row in the source data and includes the following fields: prompt (user instruction), system_prompt (system instruction, separated), and optional tools (structured description of available tools, when the source row defines tools). The data is stored in CSV format, with nested values encoded in JSON. The dataset also provides two auxiliary files, summary.md and null_or_empty_rows.md, for statistical summaries (such as the number of extracted rows and failed rows) and recording row indices where extraction results are null or empty prompts, respectively. This dataset is generated by the Nemotron Post-Training v3 prompt extractor workflow and aims to provide high-quality prompt data for instruction following and calendar-related tasks in reinforcement learning.





