Nemotron-RL-Safety-v1-prompt-only
收藏资源简介:
Nemotron-RL-Safety-v1-prompt-only数据集是从源数据集nvidia/Nemotron-RL-Safety-v1中专门提取的仅包含提示(prompt-only)的子集。该数据集旨在提供经过整理的提示文本,适用于需要纯提示数据进行模型微调、评估或安全对齐研究的场景。数据集核心文件为prompts.csv,其中每条记录包含提取的prompt文本、分离的system_prompt以及当源数据定义可用工具时的结构化tools信息(嵌套值以JSON格式编码在CSV单元格内)。此外,数据集还包含summary.md(统计摘要,包括源行计数、提取行计数、计数差异和失败提示计数)和null_or_empty_rows.md(记录提示提取结果为空的索引)。数据集规模为89066条有效提取行,2条失败提示行,总行数较源数据集减少2条。该数据集通过Nemotron Post-Training v3提示提取器工作流生成并上传。
The Nemotron-RL-Safety-v1-prompt-only dataset is a prompt-only subset specially extracted from the source dataset nvidia/Nemotron-RL-Safety-v1. This dataset aims to provide curated prompt texts, suitable for scenarios where pure prompt data is required for model fine-tuning, evaluation, or safety alignment research. The core file of the dataset is prompts.csv, where each record contains the extracted prompt text, the separated system_prompt, and structured tools information when the source data defines available tools, with nested values encoded in JSON format within the CSV cells. In addition, the dataset also includes summary.md, a statistical summary covering source row count, extracted row count, count difference and failed prompt count, and null_or_empty_rows.md, which records the indices of prompts with empty or null extraction results. The dataset has 89,066 valid extracted rows and 2 failed prompt rows, with the total number of rows being 2 less than that of the source dataset. This dataset was generated and uploaded via the Nemotron Post-Training v3 Prompt Extractor workflow.




