MemoryAsModality/swebench-verified-kimi-k2p6-traces
收藏资源简介:
该数据集包含在princeton-nlp/SWE-bench_Verified上使用fireworks_ai/kimi-k2p6-high模型和基于mini-swe-agent的工具生成的推理轨迹,旨在用于软件工程智能体的研究和蒸馏。它提供三种配置:raw_trajectories(每SWE-bench实例一行,包含补丁、清理结果JSON、完整轨迹JSON、消息JSON和评估报告JSON,如果可用)、turns(每对话轮次一行,包括助理的reasoning_content,如果可用)和manifest(每实例的紧凑元数据,用于过滤和统计)。数据集包含488个实例、51939个轮次,官方评估完成477个,其中334个已解决,143个未解决,解决率在已完成评估中为70.0%。
This dataset contains reasoning traces generated on princeton-nlp/SWE-bench_Verified using fireworks_ai/kimi-k2p6-high with a mini-swe-agent based harness. It is intended for research and distillation of software-engineering agents. It includes three configs: raw_trajectories (one row per SWE-bench instance with the patch, sanitized result JSON, full trajectory JSON, message JSON, and eval report JSON where available), turns (one row per conversational turn, including assistant reasoning_content when available), and manifest (compact per-instance metadata for filtering and accounting). The dataset has 488 instances, 51939 turns, with 477 official evals completed, 334 resolved, 143 unresolved, and a resolved rate of 70.0% over completed evals.



