Nemotron-SFT-CUDA-v1-prompt-only
收藏资源简介:
Nemotron-SFT-CUDA-v1-prompt-only是一个从源数据集nvidia/Nemotron-SFT-CUDA-v1中提取的仅包含提示部分的数据集。核心文件prompts.csv包含提取出的用户提示(prompt)、分离的系统提示(system_prompt),以及当源数据定义可用工具时的结构化工具描述(tools),嵌套值以JSON格式在CSV单元格内编码。数据集规模为2276条提示记录,提取成功且行数一致,无失败记录,并提供总结文件(summary.md)和空值行索引文件(null_or_empty_rows.md)用于统计和异常记录。该数据集适用于大语言模型的指令微调(SFT)、提示工程分析、系统提示与用户提示分离研究,以及工具调用场景下的结构化提示处理等任务,由Nemotron Post-Training v3提示提取器工作流生成并上传。
Nemotron-SFT-CUDA-v1-prompt-only is a dataset specifically extracted from the source dataset nvidia/Nemotron-SFT-CUDA-v1, containing only the prompt portions. The core file is prompts.csv, where each record corresponds to a row in the source data and includes extracted user prompts, separated system prompts, and structured tool descriptions when tools are defined in the source data. Nested values are encoded in JSON format within CSV cells. The dataset scale consists of 2276 successfully extracted prompt records, with no extraction failures, and the row count remains consistent before and after extraction. Additionally, summary files (summary.md) and null or empty row index files (null_or_empty_rows.md) are provided to record statistical information and anomalies from the extraction process. This dataset is suitable for tasks such as instruction fine-tuning (SFT) for large language models, prompt engineering analysis, research on separating system and user prompts, and structured prompt handling in tool invocation scenarios. It is generated and uploaded by the Nemotron Post-Training v3 prompt extractor workflow.




