jkminder/model-raising-reflection-end-eval
收藏资源简介:
这是一个用于评估的保留集,专门针对基于宪章指导的预训练反思,放置在文档末尾(reflection_end)。每个数据行包含一个dolma3网络文档以及配对的第一人称和第三人称反思,这些反思引用了文档实质性涉及的宪章部分([X.Y])。数据集使用冻结的生产管道(Qwen3.5-35B-A3B-FP8模型,提示generator_reflection_v7.md,宪章为ModelRaisingConstitution v0.2)生成,以确保黄金标准与训练标签的生产方式一致。Qwen作为参考生成器;此集用于评估其他模型与其对比。数据集构建确保与训练数据零重叠,来源为allenai/dolma3_mix-6T,并采用互补分片策略。它包含有害和良性内容的均衡样本(各50%),以镜像训练注释流,其中有害内容基于安全分类器得分(safety_score >= 3)定义。数据集列包括文档ID、源文本、安全得分、反思文本、位置信息等,并提供了行数、有害/良性比例、引用比例等统计信息。注意事项指出安全分类器在阈值上存在低精确度问题,建议用户谨慎解释有害标签。
A held-out evaluation set for charter-guided pretraining reflections, placed at the document end (reflection_end). Each row is one dolma3 web document plus a paired first-person / third-person reflection that cites charter sections ([X.Y]) where the document substantively engages with them. Generated with the frozen production pipeline (Qwen3.5-35B-A3B-FP8, prompt generator_reflection_v7.md, charter ModelRaisingConstitution v0.2) so the gold matches how the training labels were produced. Qwen is the reference generator; this set is for evaluating other models against it. The dataset is built to have zero overlap with training data, sourced from allenai/dolma3_mix-6T using a complementary shard strategy. It includes balanced samples of harmful and benign content (50/50) to mirror the training annotated stream, with harmful defined by safety classifier score (safety_score >= 3). Columns include doc_id, source text, safety score, reflection texts, position details, etc., with statistics on row count, harmful/benign ratio, citation rate, etc. Caveats note low precision of the safety classifier at the threshold, advising users to interpret harmful labels with caution.




