遇见数据集

DefBench v2 — Calibrated-Refusal Benchmark (30 paired items) with Sovereign-Fleet and Cross-Lab Frontier Runs

收藏
Zenodo2026-08-14 更新2026-08-20 收录
官方服务:

资源简介:

DefBench v2 expands the calibrated-refusal instrument from 14 to 30 items: 15 harmful requests each paired with a benign twin sharing surface vocabulary, so keyword-matching scores near chance and over-refusal is a measured failure, not a safe default. Deterministic grading (REFUSE/COMPLY regex, temperature 0); degenerate constant strategies score exactly 50%. Contents: defbench_v2.py (bank + runner), defbench_v2_run_2026-08-14.json (5-model sovereign-fleet run, finding: refusal-trained lanes over-refuse — sov6-ethics 66.7% calibration, sov-candidate 76.7%; sov6-preservation 100%), defbench_crosslab_2026-08-14.json (7 frontier labs via OpenRouter, all 100% calibrated; qwen3.7-plus unreachable, recorded not zero-filled). Part of the CSOAI GSPC instrument family. Signed measurement cards for these runs are publicly verifiable at https://csoai.org/verify.

提供机构:
Zenodo
创建时间:
2026-08-14
二维码
社区交流群
二维码
科研交流群
商业服务