jkminder/model-raising-pbsft-safety-180k
收藏资源简介:
model-raising-pbsft-safety-180k 是一个包含182,688个安全相关提示的“章程感知配对SFT数据集”。每一行将一个用户提示与针对同一提示的三种助手回复配对:1. 一个“章程感知”回复,使用`[X.Y]`标记内联引用价值章程;2. 一个“章程不可见”的相同回复(无标记、无章程词汇);3. 提示源数据集中附带的“原始回复”。该数据集是EPFL DLAB的“模型提升”项目的一部分,旨在研究章程注释预训练和后训练之间的“角色绑定桥梁”,特别支持“安全百分比消融”(在SFT中变化安全相关数据的比例)。所有三种回复变体共享相同的用户回合,因此切换模型训练内容只需单行列交换。数据生成基于两个安全/越狱领域数据集(WildJailbreak和WildGuardMix),使用Qwen3.5-35B-A3B-FP8模型和ModelRaisingConstitution v0.2章程生成回复。数据集仅包含安全相关提示(包括有害和良性提示),排除了无原始回复的行和HarmfulQA数据。
A charter-aware paired SFT dataset of 182,688 safety-relevant prompts. Each row pairs a user prompt with three assistant responses to the same prompt: 1. a charter-aware response that cites a value constitution inline with `[X.Y]` markers, 2. a charter-invisible rendering of that same response (no markers, no charter vocabulary), and 3. the original response that shipped with the prompts source dataset. It is part of the Model Raising project (EPFL DLAB), built to study the persona-binding bridge between charter-annotated pretraining and post-training, and specifically to support a safety-percentage ablation (varying the fraction of safety-relevant data in SFT). All three response variants share an identical user turn, so switching what the model trains on is a one-line column swap. Prompts are drawn from two safety/jailbreak-domain datasets (WildJailbreak and WildGuardMix), and responses were generated by Qwen3.5-35B-A3B-FP8 prompted with the value constitution. The dataset includes both harmful and benign prompts but excludes rows without original responses and HarmfulQA data.




