2026-08-13-difficult-advice-v2
收藏资源简介:
该数据集是“synth difficult_advice run”实验的阶段性快照,支持可恢复生成缓存。数据集包含多个阶段的JSONL文件(命名格式为stage_<n>_<name>.jsonl)以及一个manifest.json清单文件。生成日期为20260813_233520。使用的宪法文件为constitutions/claude_distilled_12_principles_mid/constitution.md,源代码仓库为https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git(特定提交哈希)。模型为每个阶段使用的模型(具体信息见manifest.json),生成配置也包含在manifest.json中。数据集通过命令“uv run synth run --config configs/data/synth/difficult_advice.yaml”生成。该数据集可用于研究基于宪法指导的合成数据生成,特别是针对困难建议场景的多阶段生成过程。
This dataset is a checkpoint of the synth difficult_advice run experiment, supporting resumable generation cache. It contains multiple stage JSONL files (named in the format stage_<n>_<name>.jsonl) and a manifest.json file. The generation date is 20260813_233520. The constitution file used is constitutions/claude_distilled_12_principles_mid/constitution.md, and the source code repository is https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git (specific commit hash). The model used for each stage (details in manifest.json) and generation configuration are also included in manifest.json. The dataset was generated by the command uv run synth run --config configs/data/synth/difficult_advice.yaml. This dataset can be used to study constitution-guided synthetic data generation, especially the multi-stage generation process for difficult advice scenarios.
数据集概述:2026-08-13-difficult-advice-v2
基本信息
- 许可证:MIT
- 语言:英语
- 标签:constitutional-ai、synthetic-data、sft
- 生成日期:2026-08-14
- 数据集大小:1,968 条训练记录(计划 2,000 条)
数据集用途与设计
- 实验名称:Difficult-advice v2,采用 Teaching-Claude-Why 方案,修复了 v1 的四个缺陷(强制场景多样性、去重门控、语音检查、元数据覆盖修复、股票开场审计),并启用宪法提示缓存。
- 宪法依据:使用
claude_distilled_12_principles_mid宪法(与冻结的九个原则快照claude_distilled_09_principles_mid_20260804字节一致),源文件位于 GitHub 仓库,确保与模型评估模型语料库共享同一对齐目标。
生成配置
- 模型:
anthropic/claude-haiku-4.5(场景、提示草稿、回复草稿)和anthropic/claude-sonnet-5(提示精炼、回复重写),通过 OpenRouter 调用;运行中段后固定使用 Anthropic 第一方路由(禁用回退)。 - 温度参数:场景 1.1、草稿 1.0、精炼 0.7、回复 1.0、重写 0.7;各阶段 max_tokens 2048–12288。
- 推理:所有阶段均禁用隐藏推理;精炼和重写阶段使用宪法缓存断点。
- 成本:约 164 美元(计量费用)。
数据结构与模式
- 主要文件:
stage_8_export_sft.jsonl,每行一条训练记录,包含:messages:系统、用户、助手(含content和reasoning_content)metadata:场景 ID、特征 ID/名称/文本、领域、情境、快捷方式、分块来源(块 ID、粒度、分组策略、块数)
- 另有
stage_N_*.jsonl阶段性快照文件(用于断点恢复)和manifest.json(记录 git SHA、各阶段使用量及消融信息)。
数据缺失与质控
- 缺失 32 条:18 条被 Anthropic 第一方及所有第三方主机拒绝(真实拒绝率约 0.9%);14 条反复违反语音契约检查(宪法诵读/股票开场)或重写标签格式。通过顶充流程重新采样,恢复 16 条记录。
- 质量检查:
- 场景级:2,000 个场景中 0 个重复(0.86 嵌入余弦),每个性状的 top 8-gram 占 1.3%,distinct-2 为 0.64,v1 的 46.9% top-10 领域集中问题在生成时已修复。
- 语料级(1,952 条):通过检查,0 个严重问题,1 个警告——模式扫描分类器对其自身顶级发现("拒绝机制非目标"修辞形态,报告 99.7% 广泛存在)未能通过召回率健全性检查,该数字不可靠。
- 重要提示:v1 镜像
synthdoc-v2-difficult-advice是 v2 前所有结果的记录,请勿混合使用两个版本。





