SATD Repayment Dataset
收藏资源简介:
该数据集由加拿大女王大学、曼尼托巴大学和北卡罗来纳州立大学的研究团队创建,旨在支持自动化技术债务(SATD)还款的研究。数据集包含Python和Java两种编程语言的SATD还款样本,总计156,069条,其中Python数据集58,722条,Java数据集97,347条。数据来源于多个开源代码库的提交历史,经过10步过滤流程提取,并使用LLM进行相关性过滤,确保数据的高质量和代表性。数据集的应用领域主要集中于软件工程中的技术债务管理,旨在通过自动化手段解决代码中的技术债务问题,提升代码质量和维护效率。
This dataset was created by research teams from Queen's University (Canada), the University of Manitoba, and North Carolina State University, with the objective of supporting research on automated Self-Admitted Technical Debt (SATD) repayment. It encompasses a total of 156,069 SATD repayment samples across two programming languages: Python and Java, including 58,722 samples for the Python subset and 97,347 samples for the Java subset. The data is extracted from the commit histories of multiple open-source code repositories through a 10-step filtering workflow, followed by relevance filtering using Large Language Models (LLMs) to guarantee high data quality and representativeness. The primary application domain of this dataset is technical debt management in software engineering, where it aims to address technical debt in code via automated methods to enhance code quality and maintenance efficiency.




