agency-transfer-benchmark
收藏资源简介:
Agency Transfer Benchmark(前沿草案0.2版)是一个聚合与来源数据集,作为交互式Agency Transfer Benchmark(https://miguelguerrero.eu/agencytransfer/)的配套资源。该数据集由Miguel Guerrero在剑桥ERA项目中独立研究产出,聚焦于前沿AI、有害操纵与选举安全。数据集包含以下内容:(1)2022-2026年前沿模型权威注册表,区分了至少100B总参数的开权重检查点与参数数量未公开的托管前沿API;(2)来自InfoOpsBench v2、SaferAI APE/MASK、Anthropic的只助益型智能体评估、MASK和DisElect的已发布观察结果,并附有来源链接;(3)一个历史失败的12端点APE衍生OpenRouter流水线审计,包含聚合结果、哈希、消毒路由元数据、成本和失败状态;(4)每个项目测试端点的简短研究笔记。所有记录均不包含原始有害提示或生成内容、个人数据、凭证、目标材料或当前选举操作内容。数据集中包含一个312次调用的历史失败试点(OpenRouter pilot),该试点因设计缺陷被永久排除在比较结果之外,仅保留用于可复现性和失败分析。数据集提供多个文件(JSON/CSV格式),通过Hugging Face Hub的配置视图分别展示。每个文件的train标签仅是技术容器,并非声明这些评估结果构成模型训练划分。预期用途包括:复现网站、审计来源、在精确协议或comparabilityGroup内比较原生行。明确警告:跨仪器结果不能合并为标量、排名、趋势或“Agency Transfer分数”。缺失结果不等于零。InfoOps合规性不等于人类说服力;APE检测的是说服尝试而非说服成功;MASK说谎不等于诚实的补集;模拟智能体任务完成不等于部署行为或选举效果。数据集采用CC BY 4.0许可证,代码为Apache-2.0。
The Agency Transfer Benchmark (Frontier Draft v0.2) is an aggregated and sourced dataset, accompanying the interactive Agency Transfer Benchmark (https://miguelguerrero.eu/agencytransfer/). This dataset is the independent research output of Miguel Guerrero in the Cambridge ERA project, focusing on frontier AI, harmful manipulation, and election security. The dataset includes: (1) an authoritative registry of frontier models from 2022-2026, distinguishing between open-weight checkpoints with at least 100B total parameters and hosted frontier APIs with undisclosed parameter counts; (2) published observations from InfoOpsBench v2, SaferAI APE/MASK, Anthropics only-helpful agent evaluations, MASK, and DisElect, with source links; (3) a historical failed 12-endpoint APE-derived OpenRouter pipeline audit, including aggregated results, hashes, sanitized routing metadata, costs, and failure status; (4) short research notes for each project test endpoint. All records contain no original harmful prompts or generated content, personal data, credentials, target materials, or current election manipulation content. The dataset includes a historical failed pilot (OpenRouter pilot) with 312 calls, permanently excluded from comparative results due to design flaws, retained only for reproducibility and failure analysis. The dataset provides multiple files (JSON/CSV format) displayed via Hugging Face Hubs configuration views. The train label in each file is merely a technical container and does not indicate that these evaluation results constitute training splits. Intended uses include: reproducing the website, auditing sources, comparing native rows within precise protocols or comparabilityGroup. Explicit warnings: cross-instrument results cannot be aggregated into scalars, rankings, trends, or Agency Transfer scores. Missing results are not zero. InfoOps compliance does not equal human persuasiveness; APE detects persuasion attempts, not persuasion success; MASK lying is not the complement of honesty; simulated agent task completion does not equal deployment behavior or election effects. The dataset is licensed under CC BY 4.0, and code under Apache-2.0.
Agency Transfer Benchmark 数据集概述
基本信息
- 数据集名称:Agency Transfer Benchmark(frontier draft 0.2)
- 许可证:CC BY 4.0
- 语言:英语
- 规模:n<1K(小于1000条)
- 标签:模型评估、AI安全、操纵、说服、前沿模型
- 归属:Miguel Guerrero 的 Cambridge:ERA 研究项目,属于独立研究,非官方ERA基准
数据集内容
该数据集为聚合与溯源数据集,包含以下组件:
1. 前沿模型注册表(2022–2026)
- 区分开放权重模型(≥100B总参数)与托管前沿API(参数不公开)
2. 来源链接的已发表观测数据
- InfoOpsBench v2
- SaferAI APE/MASK
- Anthropic的helpful-only智能体评估
- MASK
- DisElect
3. 历史失败审计(OpenRouter试点)
- 12端点、312次调用的管线审计
- 包含聚合结果、哈希、路由元数据、成本和失败状态
- 明确排除在比较结果之外,仅用于可复现性和失败分析
4. 研究笔记
- 每个受测端点的简短研究说明
数据安全声明:不包含原始有害提示、生成内容、个人数据、凭证、定向材料或当前选举操作内容。
文件结构与配置
数据集通过多个配置(config)暴露不同文件:
| 配置名 | 对应文件 |
|---|---|
frontier_observations |
frontier-observations.json |
frontier_models |
frontier-models.csv |
infoopsbench |
infoopsbench-2026-07-26.csv |
ape_mask_glm52 |
saferai-glm52-ape-mask.csv |
agentic_influence |
anthropic-agentic-influence.csv |
diselect |
diselect-summary.csv |
mask_original |
mask-original-results.csv |
historical_failed_pilot |
runs/2026-08-10-ape-frontier-pilot-v01/aggregate.csv |
注意:train标签仅为Hugging Face Viewer的技术容器,不代表模型训练集划分。
使用边界与测量原则
- 仅允许在精确协议或comparabilityGroup内比较原生行
- 禁止跨工具合并结果,不产生标量分数、排名或趋势
- 缺失结果不等于零值
- 各指标不可互换:
- InfoOps合规 ≠ 人工说服
- APE检测的是说服尝试,而非说服成功
- MASK说谎 ≠ 诚实度的补集
- 模拟智能体任务完成 ≠ 实际行为或选举影响
- OpenRouter试点是失败审计,不是模型比较
前沿模型纳入规则
- 开放权重:要求精确的公开指令/聊天检查点,总参数≥100B
- 托管模型:采用提供商文档中的前沿/旗舰规则,参数记为"未公开"
- 100B阈值是采样规则,非科学边界
- 总参数与激活参数不可互换
引用与许可
- 项目聚合与元数据:CC BY 4.0,受上游权利约束
- 上游材料不重新授权
- 代码:Apache-2.0
- 每次复用需同时引用本项目及适用的主要来源
- 详细文档见:PROVENANCE.md、METHODS.md、RESPONSIBLE_RELEASE.md





