Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
收藏资源简介:
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only 是一个从源数据集 nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 中专门提取提示(prompt)内容的数据集。该数据集包含 96,968 条提取记录,每条记录对应源数据集中的一行数据。主要数据文件 prompts.csv 包含三个关键字段:完整的提示文本(prompt)、分离的系统提示(system_prompt)以及当源行定义可用工具时的结构化工具描述(tools),其中嵌套值以 JSON 格式编码存储在 CSV 单元格内。数据集还包含两个辅助文件:summary.md 提供源行计数、提取行计数、计数差异和失败提示计数的统计摘要;null_or_empty_rows.md 记录提示提取结果为 null 或空值的行索引。该数据集专为提示工程、对话系统开发、工具调用场景的模型训练或评估而设计,适用于需要结构化提示和工具定义的研究与应用场景。数据集通过 Nemotron Post-Training v3 提示提取器工作流创建,属于后训练(post-training)数据整理范畴。
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only is a dataset that specifically extracts prompt content from the source dataset nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1. This dataset contains 96,968 extracted records, with each record corresponding to one row in the source dataset. The primary data file prompts.csv includes three key fields: the full prompt text (prompt), the separated system prompt (system_prompt), and the structured tool description (tools) when the source row defines available tools, where nested values are encoded in JSON format and stored within CSV cells. The dataset also includes two auxiliary files: summary.md provides a statistical summary of source row count, extracted row count, count difference, and failed prompt count; null_or_empty_rows.md records the row indices where the prompt extraction results are null or empty. This dataset is specifically designed for model training or evaluation in prompt engineering, conversational system development, and tool invocation scenarios, and is suitable for research and application scenarios requiring structured prompts and tool definitions. The dataset is created via the Nemotron Post-Training v3 prompt extractor workflow and falls under the category of post-training data curation.




