遇见数据集

PENTACET data - 23 Million Contextual Code Comments and 500,000 SATD comments

收藏
Zenodo2023-03-23 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

PENTACET is a large Curated Contextual Code Comments per Contributor and the most extensive SATD data. We mine 9,096 Open Source Software Java projects with a total of 435 million LOC. The outcome is dataset with 23 million code comments, preceding and succeeding source code context for each comment, and more than 500,000 comments labeled as SATD, including both ‘Easy to Find’ and ‘Hard to Find’ SATD.

PENTACET是一款按贡献者维度整理的上下文丰富型代码注释大型数据集,同时也是目前规模最为庞大的自述式技术债务(Self-Admitted Technical Debt,SATD)数据集。本数据集的构建过程中共采集了9096个Java语言开源软件项目,总代码行数(Lines of Code,LOC)达4.35亿。最终形成的数据集包含2300万条代码注释,每条注释均附带对应的前置与后置源代码上下文,且其中超过50万条注释被标记为自述式技术债务,涵盖"易发现"与"难发现"两类自述式技术债务注释。

提供机构:
Zenodo
创建时间:
2023-03-23
二维码
社区交流群
二维码
科研交流群
商业服务