AweAI-Team/Scale-SWE-Distilled-DeepSeek-v4-Pro-High-41k
收藏资源简介:
Scale-SWE数据集是一个大规模的软件工程任务数据集,来源于GitHub上的600多万个pull请求和23000多个仓库,覆盖了5200个不同的仓库。它包含10万个高质量实例,以及从DeepSeek v3.2模型提取的7.1万个轨迹(总计3.5B token)。数据集专注于Python编程语言,每个实例包括唯一标识符、仓库信息、工作目录、问题描述、补丁文件、测试用例等字段,旨在训练和评估编码代理在真实可执行任务上的性能。
The Scale-SWE dataset is a large-scale software engineering task dataset sourced from over 6 million pull requests and 23,000+ repositories on GitHub, covering 5,200 repositories. It contains 100,000 high-quality instances and 71,000 trajectories distilled from DeepSeek v3.2 (totaling 3.5B tokens). The dataset focuses on the Python programming language, with each instance including fields such as a unique identifier, repository information, working directory, problem statement, patch files, test cases, etc., aiming to train and evaluate coding agents on real executable tasks.




