jkminder/model-raising-pbsft-eval
收藏资源简介:
该数据集是一个保留的评估集,专门用于基于宪章(charter)的配对监督微调(SFT)。每一行数据包含一个用户提示和两个助手回复版本:带引用版本(带有[X.Y]宪章标记,表示助手在回复中附加了所依据的宪章部分)和不带引用版本(移除了所有标记和宪章词汇的相同回复)。宪章采用ModelRaisingConstitution v0.2,涵盖6个领域的35个价值元素,每个元素以[X.Y]标识。数据集使用冻结的生产管道生成,基于Qwen3.5-35B-A3B-FP8模型和v11提示,确保黄金回复与训练标签的生成方式一致,旨在评估其他模型相对于该参考模型的性能。数据集包含9,993行,列包括来源(如wildjailbreak、wildguardmix、wildchat)、原始标识符、带引用和不带引用的消息列表、是否包含引用、引用的宪章元素列表、伤害类别(如有害、良性等)和元数据。此外,还提供了经过Claude清理的列,用于纠正原始生成中的引用错误(如过度引用、错误部分标识),以及有效引用列,定义了每个提示可接受的宪章引用集合,用于更全面地评估模型对宪章的理解。统计信息显示,63%的行包含引用,伤害类别分布为对抗性有害3,963行、对抗性良性2,427行等,来源以WildJailbreak为主。引用分布显示,清理后引用标记总数为7,075个,集中在伤害与安全、尊严与权利、诚实等领域。
A held-out evaluation set for charter-aware paired supervised fine-tuning (SFT). Each row contains one user prompt with two assistant renderings of the same response: a cited version (with [X.Y] charter markers, indicating the assistant attaches the constitution section it is acting on) and an uncited version (the same response with the markers and any charter vocabulary removed). The charter is the ModelRaisingConstitution v0.2, which includes 35 value elements across 6 domains, each addressed as [X.Y]. The dataset is generated using a frozen production pipeline based on the Qwen3.5-35B-A3B-FP8 model and prompt v11, ensuring that the gold responses match how the training labels were produced, and is intended for evaluating other models against this reference. It consists of 9,993 rows with columns including source (e.g., wildjailbreak, wildguardmix, wildchat), source identifier, messages with and without citations, whether citations are present, list of cited charter elements, harm category (e.g., harmful, benign), and metadata. Additionally, it provides Claude-cleaned columns to correct citation errors in the original generation (such as over-citation or wrong section IDs) and valid-citation columns that define the acceptable set of charter citations for each prompt, enabling a broader assessment of model understanding. Statistics show that 63% of rows contain citations, with harm category distribution including adversarial_harmful (3,963 rows), adversarial_benign (2,427 rows), etc., and sources dominated by WildJailbreak. Citation distribution indicates 7,075 cleaned markers, concentrated in domains like Harm & Safety, Dignity & Rights, and Honesty.




