parsaidp/swe-bench-verified-raw-traces-qwen3-coder
收藏资源简介:
该数据集名为SWE-bench Verified raw mini-SWE-agent traces,包含来自20250802_mini-v1.0.0_qwen3-coder-480b-a35b-instruct的原始mini-SWE-agent轨迹,专门用于SWE-bench Verified任务。数据集提供多种配置:raw配置包含原始轨迹数据,其中easy分割使用来自parsaidp/SWE-bench_Verified_easy的194个实例ID,但轨迹内容是本地的mini-SWE-agent/Qwen跟踪;bridge_full和bridge_easy配置为Megatron Bridge格式的数据,适用于所有500个轨迹或仅194个easy轨迹,并分为训练、验证和测试集;此外,还有内存条件配置(如raw_memo和bridge_easy_memo),这些配置对easy子集进行了重写,将纯探索性shell调用替换为记忆回忆,并保留非探索性调用。数据集文件包括JSONL格式的轨迹数据、元数据文件(如实例ID列表和源信息),旨在支持软件工程基准测试和代理轨迹分析。
The dataset is named SWE-bench Verified raw mini-SWE-agent traces, containing raw mini-SWE-agent trajectories from 20250802_mini-v1.0.0_qwen3-coder-480b-a35b-instruct for SWE-bench Verified. It offers multiple configurations: the raw configuration includes raw trajectory data, with the easy split using exactly 194 instance IDs from parsaidp/SWE-bench_Verified_easy but with local mini-SWE-agent/Qwen traces; the bridge_full and bridge_easy configurations provide data in Megatron Bridge format for all 500 traces or only the 194 easy traces, split into training, validation, and test sets; additionally, there are memory-conditioned configs (e.g., raw_memo and bridge_easy_memo) that rewrite the easy subset by replacing purely exploratory shell calls with memory recollections and retaining non-exploratory calls. Dataset files include trajectory data in JSONL format, metadata files (such as instance ID lists and source information), and are designed to support software engineering benchmarking and agent trajectory analysis.




