CodeTrace-Agent Benchmark Lifecycle Reliability Research Compendium
收藏资源简介:
This compendium supports Auditing the Lifecycle Reliability of a Repository-Level Coding-Agent Benchmark: A Multi-Checkpoint Empirical Study. Version 1.0.3 repairs public reproduction entry points and adds fixed-revision task restoration manifests and sanitized stage evidence. The unchanged frozen primary analysis conditions on reference-control passage and registered-verifier acceptance of the variant: 14/14 concerns and 17/17 controls escaped (risk difference 0.0 percentage points; Newcombe 95% CI -21.5 to 18.4). An explicitly post-execution descriptive sensitivity includes all reference-passing designs: 14/15 versus 17/19 (risk difference +3.86 percentage points; 95% CI -20.50 to 25.43). These tailored constructions do not establish screening-label discrimination or defect prevalence. Public statistical recomputation requires only Python's standard library and included tables. A new Docker execution additionally requires the fixed upstream data and compatible container images; this release does not claim a new clean-environment Docker replication. Reviewer identities, private sheets and keys, credentials, copied workspaces, raw model dialogues, machine-local logs, and manuscripts are excluded. Team-created data and documentation are CC BY 4.0; analysis code is MIT licensed.



