遇见数据集

HandoffBench: An Evaluation Suite and Dataset for Reasoner-Verifier Interfaces in LLM Research Validation

收藏
Zenodo2026-04-20 更新2026-05-26 收录
官方服务:

资源简介:

We release HandoffBench, an evaluation suite, protocol, and run-level dataset for measuring verification behaviour in LLM-based research-validation pipelines. The suite targets a question that existing LLM-as-judge benchmarks do not isolate: does a reasoner–verifier architecture actually perform evidence-grounded checking, or does it produce fluent review text that merely resembles checking? HandoffBench consists of (i) 6 adversarially constructed research-validation tasks across four domains (NLP, statistics, biology, ethics), each embedding at least one known counter-evidence target and at least one uncertainty; (ii) 13 architectural conditions—a pre-registered three-way (C0 monolithic / C1 structured-handoff / C2 unstructured-split), a post-hoc prose-handoff (C1.5), and a nine-way field-level ablation of the C1 packet (C1-KC / C1-CL / C1-CR / C1-EV / C1-UN / C1-AC / C1-SC / C1-CORE / C1-NONE); (iii) a three-judge LLM evaluation protocol (GPT-4o + Claude Sonnet 4 + Gemini 2.5 Pro) with a four-metric rubric (Degradation Rate, Counter-Evidence Rate, Criterion Alignment Rate, Referent Continuity Score) and a 2-judge sensitivity re-aggregator; (iv) a JSON handoff packet schema with nine required fields and an auditable ablation mask pipeline; and (v) the full run-level corpus: 312 reasoner–verifier trajectories × 3 judges = 936 per-judgment JSON records, each including full reasoner output, verifier report, handoff packet, verifier-facing JSON, and per-metric judge rationales. We validate HandoffBench by demonstrating its discriminative and diagnostic power. On the pre-registered three-way design, C0 and C2 are statistically indistinguishable on all four metrics (all pairwise p > 0.5), while C1 separates sharply (DR 0.833 → 0.056, CER 0 → 0.611, CAR 0.60 → 0.97, all p < 0.001). The field-level ablation further separates a sufficient-core subset (C1-CORE, three fields: claims_to_verify + criteria + known_conflicts) that is statistically identical to full C1, and a total-collapse condition (C1-NONE, empty packet contents but intact keys: DR 0.958, CER 0.042, CAR 0.00). Inter-judge agreement on the full 312-run corpus is moderate-to-substantial (Fleiss' κ = 0.584 DR, 0.587 CER; ICC(2,1) = 0.761 RCS, 0.793 CAR), and a post-hoc 2-judge re-aggregation (dropping Claude Sonnet 4 as judge) preserves the main C1 separation. A follow-up with Claude Opus 4.7 (Constitutional-AI-aligned) shows the benchmark captures tuning-specific (DR, CAR) vs. tuning-invariant (CER) failure modes. We argue these results together establish HandoffBench as a compact, reusable evaluation object for reasoner–verifier interface design, orthogonal to existing end-task benchmarks. The artifact (code, tasks, schemas, all 312 trajectories, all 936 judge records, Croissant metadata, datasheet) is released under permissive licenses; see §9.

提供机构:
Zenodo
创建时间:
2026-04-20
二维码
社区交流群
二维码
科研交流群
商业服务