Nemotron-RL-Safety-v1
收藏资源简介:
Nemotron-RL-Safety-v1 数据集旨在为训练奖励模型提供必要的标记比较,以区分安全、有帮助的响应和不受欢迎、不合规的输出。该数据集包含:1. 混合(开源和合成生成)的提示集合,旨在引发不同的模型漏洞;2. 安全偏好对:每个提示与一个被选中的响应和一个被拒绝的响应相关联,为奖励模型提供清晰的训练信号。被选中的响应是安全、有帮助且符合模型行为指南的,而被拒绝的响应则是不安全或不符合推荐响应策略的。数据集适用于商业用途,包含多个底层子集,涵盖内容安全风险、越狱攻击、过度拒绝、人口统计偏见和敏感内容泄露等方面。数据集采用 JSONL 格式,包含 44,941 个独特提示和 89,882 个偏好对,总磁盘大小约为 200MB。该数据集适用于强化学习通过人类反馈(RLHF)的模型对齐,以提高安全性和安全性。
The Nemotron-RL-Safety-v1 dataset is designed to provide necessary labeled comparisons for training reward models, to differentiate between safe, helpful responses and undesirable, non-compliant outputs. This dataset includes: 1. A mixed set of prompts (both open-source and synthetically generated) intended to elicit various model vulnerabilities; 2. Safe preference pairs: each prompt is associated with one selected response and one rejected response, providing clear training signals for reward models. The selected responses are safe, helpful, and compliant with the model's behavior guidelines, while the rejected responses are unsafe or violate the recommended response strategies. This dataset is available for commercial use, and comprises multiple underlying subsets covering content safety risks, jailbreaking attacks, over-refusal, demographic bias, sensitive content leakage, and other relevant areas. The dataset is stored in JSONL format, containing 44,941 unique prompts and 89,882 preference pairs, with a total disk size of approximately 200 MB. This dataset is suitable for model alignment via Reinforcement Learning from Human Feedback (RLHF) to enhance safety and security.



