DefBench v2 — Calibrated-Refusal Benchmark (30 paired items) with Sovereign-Fleet and Cross-Lab Frontier Runs
收藏资源简介:
DefBench v2 expands the calibrated-refusal instrument from 14 to 30 items: 15 harmful requests each paired with a benign twin sharing surface vocabulary, so keyword-matching scores near chance and over-refusal is a measured failure, not a safe default. Deterministic grading (REFUSE/COMPLY regex, temperature 0); degenerate constant strategies score exactly 50%. Contents: defbench_v2.py (bank + runner), defbench_v2_run_2026-08-14.json (5-model sovereign-fleet run, finding: refusal-trained lanes over-refuse — sov6-ethics 66.7% calibration, sov-candidate 76.7%; sov6-preservation 100%), defbench_crosslab_2026-08-14.json (7 frontier labs via OpenRouter, all 100% calibrated; qwen3.7-plus unreachable, recorded not zero-filled). Part of the CSOAI GSPC instrument family. Signed measurement cards for these runs are publicly verifiable at https://csoai.org/verify.



