遇见数据集

Software agent benchmark traces for repository command selection

收藏
Zenodo2026-06-16 更新2026-06-17 收录
官方服务:

资源简介:

This dataset contains benchmark traces for bounded command selection in software-agent repository debugging settings. It includes public repository traces, isolated dynamic pytest execution records, candidate command sets, receptor-state fields, split manifests, leakage and data-health audits, and final technical-validation reports. The records are intended for reusable evaluation of structured software-agent interfaces and command-selection policies. The dataset does not claim production autonomy, unrestricted shell use, open-ended software repair, or general coding-agent superiority. Version 2 corrects release metadata, binds the embedded software snapshot to public ReflexLM v0.1.2 commit 3ee693c3773c9796491babb06eb81571486ab68b, and replaces machine-specific repository-root prefixes in text provenance fields with ${REFLEXLM_ROOT}. The seven core benchmark JSONL files and reported scientific values are unchanged from version 1.

提供机构:
Zenodo
创建时间:
2026-06-15
二维码
社区交流群
二维码
科研交流群
商业服务