遇见数据集

Reproducibility Artifact for Pre-Execution Auditing for Tool-Using AI Agents: A Reproducible Mixed Benchmark for Agentic Software Engineering

收藏
Zenodo2026-08-15 更新2026-08-20 收录
官方服务:

资源简介:

This reproducibility artifact accompanies the paper Pre-Execution Auditing for Tool-Using AI Agents: A Reproducible Mixed Benchmark for Agentic Software Engineering. It contains the locked benchmark data, audit rubric, generated method outputs, aggregate results, paired comparisons, bootstrap confidence intervals, availability analysis, source-authority records, trace-replay results, benchmark runners, validators, verification records, and SHA-256 manifest supporting the manuscript. The benchmark comprises 96 externally grounded tool-action scenarios, including 72 unsafe intervention cases and 24 benign allow-control cases, evaluated across 14 audit policies. It contains 1,344 scored method-case rows in the external benchmark and a matched 1,344-row trace-replay evaluation. The artifact supports reproduction and audit of the reported safety-availability tradeoff in pre-execution control for tool-using AI agents. It preserves the study's claim boundary: the evaluation is a reproducible benchmark and controlled trace replay, not a live deployment study and not an independently human-annotated evaluation.

提供机构:
Zenodo
创建时间:
2026-08-15
二维码
社区交流群
二维码
科研交流群
商业服务