Reproducibility Artifact for Pre-Execution Auditing for Tool-Using AI Agents: A Reproducible Mixed Benchmark for Agentic Software Engineering
收藏资源简介:
This reproducibility artifact accompanies the paper Pre-Execution Auditing for Tool-Using AI Agents: A Reproducible Mixed Benchmark for Agentic Software Engineering. It contains the locked benchmark data, audit rubric, generated method outputs, aggregate results, paired comparisons, bootstrap confidence intervals, availability analysis, source-authority records, trace-replay results, benchmark runners, validators, verification records, and SHA-256 manifest supporting the manuscript. The benchmark comprises 96 externally grounded tool-action scenarios, including 72 unsafe intervention cases and 24 benign allow-control cases, evaluated across 14 audit policies. It contains 1,344 scored method-case rows in the external benchmark and a matched 1,344-row trace-replay evaluation. The artifact supports reproduction and audit of the reported safety-availability tradeoff in pre-execution control for tool-using AI agents. It preserves the study's claim boundary: the evaluation is a reproducible benchmark and controlled trace replay, not a live deployment study and not an independently human-annotated evaluation.



