SATDAUG
收藏资源简介:
SATDAUG数据集是由伯努利研究所和格罗宁根大学创建的,旨在通过平衡和增强现有数据集来提高自我承认的技术债务(SATD)检测的准确性。该数据集包含来自源代码注释、问题跟踪器、拉取请求和提交消息的多种软件开发工件,总计约2370万条记录。创建过程中,研究者使用了基于ChatGPT的语言模型进行文本增强,确保数据集的平衡性和丰富性。SATDAUG数据集主要应用于机器学习和深度学习模型的训练和评估,以解决SATD识别和分类中的类别不平衡问题,从而提高软件维护的效率和质量。
SATDAUG dataset was developed by the Bernoulli Institute and the University of Groningen, with the goal of enhancing the accuracy of Self-Admitted Technical Debt (SATD) detection by balancing and augmenting existing datasets. This dataset encompasses a diverse range of software development artifacts sourced from source code comments, issue trackers, pull requests, and commit messages, totaling approximately 23.7 million records. During the dataset construction process, researchers utilized ChatGPT-based language models for text augmentation to ensure the dataset's balance and richness. The SATDAUG dataset is primarily intended for training and evaluating machine learning and deep learning models, aiming to address the class imbalance issue in SATD identification and classification, thereby improving the efficiency and quality of software maintenance.




