DeepSWE-Gym2
收藏资源简介:
该数据集是多个高质量软件工程(SWE)数据集的合并版本,经过过滤和去重处理,旨在提升DeepSWE风格问题、基准测试以及通用编码技能的表现。数据集专门筛选了原始数据集中包含复杂/长代码问题的样本,平均样本大小为214.19KB,总未压缩大小为17.56GB,共计85974个样例。数据集由MoreThought策划并共享,采用MIT许可证。数据来源包括DeepSWE-Gym、SWE-Gym、SWE-rebench-V2-Filtered-Verified、SWE-bench-extra、SWE-Next、Multi-SWE-RL-Verified和SWE-Synth等HuggingFace数据集,以及多篇相关论文。该数据集适用于改进基准测试结果、提升通用编程能力、训练软件工程/编码代理以及改进长上下文编码任务。注意:为避免样例重叠,请勿将该数据集与其他变体或版本一起使用。
This dataset is a merged version of multiple high-quality software engineering (SWE) datasets, filtered and deduplicated to enhance DeepSWE-style problems, benchmarks, and general coding skills. The dataset specifically selects samples containing complex/long code problems from the original datasets, with an average sample size of 214.19KB, a total uncompressed size of 17.56GB, and a total of 85,974 samples. The dataset is curated and shared by MoreThought under the MIT license. Data sources include HuggingFace datasets such as DeepSWE-Gym, SWE-Gym, SWE-rebench-V2-Filtered-Verified, SWE-bench-extra, SWE-Next, Multi-SWE-RL-Verified, and SWE-Synth, as well as multiple related papers. This dataset is suitable for improving benchmark results, enhancing general programming abilities, training software engineering/coding agents, and improving long-context coding tasks. Note: To avoid sample overlap, do not use this dataset with other variants or versions.
数据集概述
DeepSWE-Gym2 是一个经过筛选和去重处理的高质量软件工程(SWE)数据集合并版本,旨在提升 DeepSWE 风格问题、基准测试及通用编码能力上的表现。
基本信息
- 名称: DeepSWE-Gym2
- 维护方: MoreThought
- 许可证: MIT
- 规模: 85,974 条样本(10K < n < 100K)
- 数据大小: 未压缩总大小约 17.56 GB,平均每条样本约 214.19 KB(部分样本可达 28 MB)
- 任务类型: 文本生成、问答、图像-文本到文本
- 语言: 英语
- 编程语言覆盖: Python、JavaScript、TypeScript、Java、C++、Rust、Go
数据来源
该数据集由以下多个 SWE 相关数据集合并、过滤和去重而成:
- MoreThought/DeepSWE-Gym
- SWE-Gym/SWE-Gym
- PrimeIntellect/SWE-rebench-V2-Filtered-Verified
- nebius/SWE-bench-extra
- TIGER-Lab/SWE-Next
- PrimeIntellect/Multi-SWE-RL-Verified
- swesynth/SWE-Synth
另参考了若干相关研究论文(详见数据集原始页面)。
用途
- 提升基准测试结果
- 增强通用编码能力
- 训练软件工程/编码代理(Agent)
- 改善长上下文编码能力
重要提示
- 避免数据重叠: 不应将此数据集与其他变体或版本混合使用,以免产生重叠样本。
- 训练资源要求高: 大多数大语言模型无法直接处理接近 28 MB 的超长样本,需使用专门脚本来进行有效训练。




