jkminder/model-raising-pbsft-instruct-300k
收藏资源简介:
一个包含30万个通用目的(WildChat)指令提示的章程感知配对SFT数据集。每行将一个用户提示与同一提示的三种助手响应配对:1)章程感知响应,使用[X.Y]标记内联引用价值章程;2)章程不可见响应(无标记,无章程词汇);3)原始WildChat-1M中随提示提供的原始响应。该数据集是EPFL DLAB的模型提升项目的一部分,专注于通用/指令领域(如协助、写作、编码、解释、角色扮演等),而非安全/越狱提示。所有三种响应变体共享相同的用户回合,因此切换模型训练内容仅需一列交换。
A charter-aware paired SFT dataset of 300,000 general-purpose (WildChat) instruct prompts. Each row pairs a user prompt with three assistant responses to the same prompt: 1) a charter-aware response that cites a value constitution inline with [X.Y] markers, 2) a charter-invisible rendering of that same response (no markers, no charter vocabulary), and 3) the original response that shipped with the prompt in WildChat-1M. It is part of the Model Raising project (EPFL DLAB). This is the general/instruct counterpart to the safety-relevant split — it holds the open-domain, in-the-wild user prompts (assistance, writing, coding, explanations, roleplay, etc.), not safety/jailbreak prompts. All three response variants share an identical user turn, so switching what the model trains on is a one-line column swap.




