govbench
收藏资源简介:
GovBench 是 GSPC(Governance, Safety, Provenance, Continuity, Conformance, Openness)评估体系中治理轴(Governance Axis)下的基准测试数据集,专注于根据欧盟 AI 法案(EU AI Act)对 AI 系统进行风险等级分类。该数据集包含 24 个样本,每个样本被标注为以下四种风险等级之一:PROHIBITED(禁止)、HIGH_RISK(高风险)、LIMITED_RISK(有限风险)、MINIMAL_RISK(最低风险)。数据集以文本分类任务的形式呈现,用于评估语言模型对欧盟 AI 法案风险层级的理解能力。数据由独立 AI 测量机构 CSOAI Ltd 发布,基于冻结的语料锚定真值进行评分。在 30 个模型(参数规模 494M 至 20B,三种架构)上的测试显示,该轴的平均难度为 0.227,区分度(最大-最小)为 0.440,有效样本数(usable_n)为 23(因一个死项被排除)。该轴尚未饱和,但已能区分模型性能。数据集附带详细的评分指南和已知限制,强调测量而非认证,并建议报告有效样本数而非原始样本数。
GovBench is a benchmark dataset under the Governance Axis of the GSPC (Governance, Safety, Provenance, Continuity, Conformance, Openness) evaluation framework, focusing on classifying AI systems into risk levels according to the EU AI Act. The dataset contains 24 samples, each labeled as one of four risk levels: PROHIBITED, HIGH_RISK, LIMITED_RISK, or MINIMAL_RISK. It is presented as a text classification task to evaluate language models understanding of the EU AI Act risk tiers. The data is released by CSOAI Ltd, an independent AI measurement organization, and scored based on frozen corpus-anchored ground truth. Tests on 30 models (parameters ranging from 494M to 20B, with three architectures) show an average difficulty of 0.227, a discrimination (max-min) of 0.440, and a usable sample count (usable_n) of 23 (due to one dead item excluded). The axis is not yet saturated but can already distinguish model performance. The dataset includes detailed scoring guidelines and known limitations, emphasizing measurement rather than certification, and recommends reporting usable sample count instead of raw sample count.
GovBench — EU AI Act 风险等级基准(GSPC 治理轴)
基本信息
- 数据集名称:GovBench — EU AI Act risk tier
- 发布机构:CSOAI Ltd(英国,公司编号 16939677),独立 AI 测量机构
- 许可协议:Apache-2.0(数据集条目),卡片为 CC-BY-4.0
- 语言:英语
- 任务类型:文本分类
- 数据规模:n < 1K(共 24 个条目)
- DOI:10.5281/zenodo.21755657
数据集内容
该基准是 GSPC 仪器六大轴之一,用于评估模型对欧盟 AI 法案风险等级的分类能力。数据条目以冻结的、基于语料锚定的真实标签为评分依据。
标签分布(共 24 条)
| 标签 | 数量 |
|---|---|
| PROHIBITED | 7 |
| HIGH_RISK | 8 |
| LIMITED_RISK | 3 |
| MINIMAL_RISK | 6 |
测量表现(30 个模型,参数量 494M 至 20B,三种架构)
| 统计量 | 值 | 说明 |
|---|---|---|
| 平均难度 | 0.227 | 正确回答某个条目的模型比例 |
| 区分度(最大−最小) | 0.440 | 最佳与最差模型间的分离程度 |
| 无效条目 | 1/24 | 所有模型均通过或均未通过,无信息量 |
| 负向区分条目 | 0 | 整体更优模型表现更差的条目(已转裁决,未删除) |
| 可用条目数 | 23 | 非无效且非负向区分的条目 |
可引用性结论:由于可用条目数(23)低于 30,本轴暂不可发布 95% 置信区间。虽然区分度 0.440 表明该轴能区分模型,但需 usable_n ≥ 30 才能将 95% Wilson 区间收窄至 ±0.169。该限制被明确公布而非引用无法支持的区间。
评分规则
- 三种结果:measured / unmeasured / failed。无法评分的生成为
UNMEASURED,不计分,也不计为错误答案。 - 报告
usable_n而非n:避免因无效条目而高估证据量。 - 区间优先于点估计:95% 区间重叠的比较标记为
NOT_RESOLVED。 - 测量而非认证:分数带仅为描述性,监管判定权归权威机构。
已知局限
- 无前沿裁判验证;评分基于精确标签或经 98.9% 准确率验证的拒绝检测器(基于 92 条人工标注响应,单一标注者,无交叉一致性)。
- 条目难度是“条目 × 模型群”的属性,而非条目独立属性;模型群偏向小模型时,难度结论描述的是该模型群。
- 条目区分度为估计值,尚未在当前模型群规模下解决。
所属体系
GovBench 是 GSPC 六大轴之一,共计 90 个条目,各轴均从实时数据集中读取计数:
| 轴 | 基准 | 任务 | 条目数 | Hugging Face | Kaggle |
|---|---|---|---|---|---|
| 治理 | GovBench | EU AI Act 风险等级 | 24 | csoai/govbench | gspc-govbench |
| 安全 | DefBench | 校准拒绝 | 14 | csoai/defbench | gspc-defbench |
| 溯源 | ProvBench | C2PA 清单存续 | 15 | csoai/provbench | gspc-provbench |
| 连续性 | PQCBench | 后量子迁移 | 13 | csoai/pqcbench | gspc-pqcbench |
| 合规性 | MCPBench | MCP 工具契约合规 | 11 | csoai/mcpbench | gspc-mcpbench |
| 开放性 | OSSBench | 许可证-用途兼容 | 13 | csoai/ossbench | gspc-ossbench |




