rh-clean-control-sft
收藏资源简介:
Clean Control SFT Mixture 是一个用于奖励黑客实验对照的干净SFT混合数据集。该数据集仅包含良性任务,不包括故意错位、易受攻击或越狱合规的数据。数据集由多种任务类型组成,包括指令跟随、数学推理、常识问答、有帮助的聊天、摘要、安全拒绝和代码纠正,共计约10,538个样本。每个样本包含messages(角色和内容的字典列表)、prompt(用户消息的平面字符串)、completion(助手消息的平面字符串)和task_type(任务类型)字段。数据集排除了不安全代码、易受攻击代码和越狱合规等类别。适用于文本生成任务,特别是需要安全、对齐模型行为的研究场景。
Clean Control SFT Mixture is a clean supervised fine-tuning (SFT) mixture dataset intended as a control for reward hacking experiments. This dataset exclusively contains benign tasks, excluding data related to intentional misalignment, vulnerable scenarios, or jailbreak compliance. It comprises multiple task categories, including instruction following, mathematical reasoning, common sense question answering, helpful chat, summarization, safe refusal, and code correction, with a total of approximately 10,538 samples. Each sample includes fields including messages (a list of dictionaries storing roles and corresponding content), prompt (a flat string of the user's message), completion (a flat string of the assistant's message), and task_type. This dataset excludes categories such as unsafe code, vulnerable code, and jailbreak compliance-related data. It is suitable for text generation tasks, particularly research scenarios requiring safe and aligned model behaviors.




