CodeTrace-Agent Benchmark Lifecycle Reliability Research Compendium
收藏资源简介:
This research compendium provides code, frozen protocols, de-identified task-level evidence, machine-readable results, an executed notebook, and editable figures for Auditing the Lifecycle Reliability of a Repository-Level Coding-Agent Benchmark: A Multi-Checkpoint Empirical Study. Version 1.0.2 adds the prospectively frozen held-out verifier challenge: 180 previously unreviewed tasks were independently screened, all 20 adjudicated concerns were repository-matched to 20 controls, and a masked designer completed 40 direct-to-parent challenges. Thirty-one were evaluable and every tailored variant escaped (14/14 concerns; 17/17 controls; risk difference 0.0 percentage points, Newcombe 95% CI -21.5 to 18.4). The result confirms constructible registered-verifier gaps but does not establish screening-label discrimination or defect prevalence. Reviewer identities, private worksheets and role keys, access tokens, copied upstream repository snapshots, raw model dialogues, per-container receipts, machine-local logs, and manuscript files are excluded. Team-created data and documentation are CC BY 4.0; analysis code is MIT licensed.



