compass-group-tue/sdf_evaluation_traits
收藏资源简介:
这些合成文档用于微调模型,以研究评估元知识——即关于AI安全评估结构特征的参数知识。文档使用Wang等人(2025)的流程生成(具体调整的提示和细节见论文)。GPT-4.1用于文档类型和想法头脑风暴,而GPT-5生成最终文档。每个.jsonl文件包含特定评估特征的文档:可验证结构、占位符、有害请求、冲突目标、不一致环境、异常访问和伦理困境。统计信息:总令牌数106M,每个特征约15M令牌。
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the pipeline of Wang et al. (2025) (see the paper for adjusted prompts and details). GPT-4.1 was used for document type and idea brainstorming, while GPT-5 generated the final documents. Each .jsonl file contains the documents for a specific evaluation trait: Verifiable Structure, Placeholder, Harmful Requests, Conflicting Goals, Inconsistent Environment, Unusual Access, Ethical Dilemma. Statistics: Total tokens: 106M, Tokens per trait: ~15M.




