BenMark
收藏资源简介:
BenMark数据集由北京理工大学等机构构建,旨在评估深度学习在识别不一致方法名称中的表现。该数据集包含1,299,186条数据,其中2,443条为不一致方法,1,296,743条为一致方法。数据集基于430个高质量项目,通过自动识别提交历史和开发者手动检查构建,减少了误报率。该数据集的应用领域为软件工程,旨在解决不一致方法名称导致的软件缺陷和维护成本问题。
The BenMark Dataset was constructed by institutions including Beijing Institute of Technology, aiming to evaluate the performance of deep learning models in identifying inconsistent method names. This dataset contains 1,299,186 entries in total, among which 2,443 are inconsistent method entries and 1,296,743 are consistent ones. It is built based on 430 high-quality software projects, through automatically identifying commit histories and manual inspection by developers, which reduces the false positive rate. Its application domain is software engineering, and it is designed to address software defects and maintenance cost issues caused by inconsistent method names.

- 1Deep Learning-Based Identification of Inconsistent Method Names: How Far Are We?北京理工大学计算机科学与技术学院, 中国电信北京研究院, 中国人民解放军军事科学院国防科技创新研究院 · 2025年



