Software agent benchmark traces for repository command selection
收藏资源简介:
This dataset contains benchmark traces for bounded command selection in software-agent repository debugging settings. It includes public repository traces, isolated dynamic pytest execution records, candidate command sets, receptor-state fields, split manifests, leakage and data-health audits, and final technical-validation reports. The records are intended for reusable evaluation of structured software-agent interfaces and command-selection policies. The dataset does not claim production autonomy, unrestricted shell use, open-ended software repair, or general coding-agent superiority. Version 2 corrects release metadata, binds the embedded software snapshot to public ReflexLM v0.1.2 commit 3ee693c3773c9796491babb06eb81571486ab68b, and replaces machine-specific repository-root prefixes in text provenance fields with ${REFLEXLM_ROOT}. The seven core benchmark JSONL files and reported scientific values are unchanged from version 1.



