遇见数据集

evavlabs/oa

收藏
Hugging Face2026-05-11 更新2026-05-31 收录
官方服务:

资源简介:

EVAV操作对齐语料库是一个用于AI部署安全审计的数据集,包含209,072个匹配对决策,这些决策来自8个前沿大语言模型(如Claude Sonnet 4、GPT-4o等),覆盖3个受监管部署领域:医疗保健预先授权、消费贷款和投资组合交易。该数据集旨在评估模型在部署现实条件下是否保留既定规则,通过匹配对因果识别方法,检测模型在压力操纵变量下的违规行为。方法学基于匹配对审计研究设计(源自Bertrand & Mullainathan 2004),采用PRNG确定性场景生成,并在8个模型×3个领域×24+种条件下应用。数据集经过机制可解释性验证——SAE探针以81.2%的准确率检测到违规状态,操纵相关特征可将违规率从100%降至0%。每个记录包含模型响应、领域、测试条件、种子、温度、配对ID、角色(基础或孪生)、决策、违规状态、失败模式和推理文本。

Each record in `corpus.jsonl` is one models response to a structured evaluation prompt under one of 24 condition types. Matched pairs share identical templates with only the targeted manipulation variable varying, enabling within-pair causal identification of violation drivers. Methodology: matched-pair audit-study design (Bertrand & Mullainathan 2004 lineage) with PRNG-deterministic scenario generation, applied across 8 frontier models × 3 domains × 24+ conditions. Validated by mechanistic interpretability — SAE probes detect the violation state at 81.2% accuracy; steering the relevant feature reduces violation rate from 100% to 0%.

提供机构:
evavlabs
二维码
社区交流群
二维码
科研交流群
商业服务