Nemotron-SWE-v1-prompt-only
收藏资源简介:
Nemotron-SWE-v1-prompt-only 是一个从源数据集 nvidia/Nemotron-SWE-v1 中提取的纯提示词数据集。该数据集旨在提供结构化的提示词内容,适用于大型语言模型的提示工程、微调或评估任务。数据以CSV文件(prompts.csv)形式组织,每条记录包含一个prompt(提示词)、一个独立的system_prompt(系统提示词),以及当源数据定义可用工具时的结构化tools(工具)信息,其中嵌套值以JSON格式编码在CSV单元格内。数据集规模为51,029个提取的提示词行,无提取失败的行,与源数据行数一致。此外,数据集还包含摘要文件(summary.md)和空值/无效行记录文件(null_or_empty_rows.md),分别提供统计信息和问题行索引。该数据集由用户jamesdborin通过Nemotron Post-Training v3提示词提取工作流生成并上传。
Nemotron-SWE-v1-prompt-only is a prompt-only dataset extracted from the source dataset nvidia/Nemotron-SWE-v1. This dataset aims to provide structured prompt content suitable for prompt engineering, fine-tuning, or evaluation tasks for large language models. The data is organized in a CSV file (prompts.csv), with each record containing a prompt, an independent system_prompt, and structured tools information when the source data defines available tools, where nested values are encoded in JSON format within CSV cells. The dataset scale is 51,029 extracted prompt rows, with no extraction failures, consistent with the source data row count. Additionally, the dataset includes summary files (summary.md) and null/empty row record files (null_or_empty_rows.md), providing statistical information and problematic row indices, respectively. This dataset was generated and uploaded by user jamesdborin through the Nemotron Post-Training v3 prompt extraction workflow.
数据集名称
- Nemotron-SWE-v1-prompt-only
数据集来源
- 原始数据集:
nvidia/Nemotron-SWE-v1
数据集内容
- 该数据集是从
nvidia/Nemotron-SWE-v1中提取的仅包含提示(prompt)的记录。 - 包含以下文件:
- prompts.csv:每条源记录对应一条提示提取记录。记录包括
prompt、分离的system_prompt以及当源行定义了可用工具时的结构化tools。嵌套值在 CSV 单元格内以 JSON 编码。 - summary.md:源行数、提取行数、行数变化量以及失败提示数。
- null_or_empty_rows.md:提示提取产生空或 null 提示的行索引。
- prompts.csv:每条源记录对应一条提示提取记录。记录包括
数据统计
- 提取行数:51,029
- 失败提示行数:0
- 行数变化量:0
标签
- nemotron
- prompt-only
- post-training
配置
- 配置名称:default
- 数据文件:训练集分割,路径为
prompts.csv
上传者
- 由
jamesdborin从 Nemotron Post-Training v3 提示提取器工作流上传。




