VmaxRL/bugpilot-bugintro-lm-modify-gpt55-repaired-34-cleaned-1k-strict-20260526T212012Z
收藏资源简介:
该数据集是一个用于软件工程语言模型修改任务的严格数据集,包含1000个样本,涵盖Go、JavaScript、Python和TypeScript四种编程语言。数据来源于34个开源仓库,涉及全栈Web产品、基础设施系统集成、富客户端UI、后端服务平台等多种软件类别。数据集通过SWE-smith工具和Azure GPT-5.5-1模型生成,并经过验证、清理、组装和去重处理,确保数据质量和唯一性。数据集特征包括代理指令、基础提交、Docker镜像、测试列表(失败到通过和通过到通过)、框架、生成器ID、黄金补丁、介绍补丁、问题陈述、仓库信息、需求、设置命令、测试命令、元数据(如测试计数、数据集ID、源仓库等)等,适用于软件缺陷修复和代码修改的机器学习训练与评估。
This dataset is a strict dataset for software engineering language model modification tasks, containing 1000 samples across four programming languages: Go, JavaScript, Python, and TypeScript. The data originates from 34 open-source repositories, covering various software categories such as full-stack web products, infrastructure systems integration, rich client UI, and backend service platforms. The dataset is generated using the SWE-smith tool and the Azure GPT-5.5-1 model, and undergoes validation, cleaning, assembly, and deduplication to ensure data quality and uniqueness. Features include agent instructions, base commit, Docker image, test lists (fail-to-pass and pass-to-pass), framework, generator ID, gold patch, introduction patch, problem statement, repository information, requirements, setup command, test command, metadata (e.g., test counts, dataset ID, source repository, etc.), and is suitable for machine learning training and evaluation in software bug fixing and code modification.




