Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
收藏资源简介:
Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only 是一个专门从源数据集 nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1 中提取提示词(prompt)的数据集。它包含9037条提取记录,每条记录对应源数据集的一行数据,以CSV文件格式组织。主要字段包括:提取的提示词(prompt)、分离的系统提示词(system_prompt)以及当源数据定义可用工具时的结构化工具信息(tools),其中嵌套值以JSON格式编码在CSV单元格内。数据集还提供了元数据文件,包括数据统计摘要和标识空或无效提示行的索引文件。该数据集是Nemotron后训练工作流的一部分,专门用于提示词提取任务,适用于指令跟随、自由形式格式化和基于提示的语言模型训练等场景。
Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only is a dataset specifically designed to extract prompts from the source dataset nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1. It contains 9,037 extracted records, each corresponding to one row of data from the source dataset, and is organized in CSV file format. Its core fields include: the extracted prompt, the separated system prompt, and structured tool information (tools) when the source data defines available tools, where nested values are encoded in JSON format within CSV cells. The dataset also provides metadata files, including a statistical summary of the data and an index file that identifies empty or invalid prompt rows. This dataset is part of the Nemotron post-training workflow, specifically tailored for prompt extraction tasks, and is applicable to scenarios such as instruction following, free-form formatting, and prompt-based language model training.




