遇见数据集

Beyond One-to-One: A Benchmark Dataset for One-to-Many Issue-Commit Links

收藏
Zenodo2025-05-28 更新2026-05-26 收录
官方服务:

资源简介:

Recovering missing links between issues and commits is crucial for effective software traceability, maintenance, and understanding how code evolves over time. While existing research has predominantly focused on recovering one-to-one issue-commit links, real-world software development often follows a one-to-many pattern, where a single issue is resolved across multiple commits. This one-to-many pattern has largely been overlooked, making it harder for existing automated methods to work effectively in real development scenarios. In this study, we introduce a large-scale dataset specifically built for one-to-many issue-commit link recovery, spanning projects written in four widely-used programming languages: Java, C++, Python, and JavaScript. To accurately assess model performance in the one-to-many setting, we propose an issue-wise evaluation strategy that moves beyond traditional link-level metrics, better reflecting the need to recover complete sets of relevant commits. We benchmark a diverse set of state-of-the-art models, including rule-based, machine learning, deep learning, and transformer-based approaches. Our results reveal that, using a balanced dataset, existing models achieve only moderate performance, with the highest F1-score reaching 73%, underscoring the need for more effective models tailored for one-to-many link recovery. We publicly release our dataset and evaluation framework to support future research in building more robust and effective models for one-to-many link recovery.

恢复问题(issue)与代码提交(commit)间的缺失关联,对于实现高效的软件可追溯性、开展软件维护工作,以及厘清代码随时间的演化规律至关重要。尽管现有研究主要聚焦于恢复一对一的问题-代码提交关联,但现实中的软件开发往往遵循一对多模式:单个问题可通过多次代码提交完成修复。这种一对多模式在很大程度上被现有研究忽视,导致现有的自动化方法难以在实际开发场景中有效运行。本研究构建了一款专为一对多问题-代码提交关联恢复任务打造的大规模数据集,其涵盖了Java、C++、Python及JavaScript四种主流编程语言编写的项目。为精准评估模型在一对多场景下的性能,我们提出了一种基于问题维度的评估策略,该策略突破了传统的关联层级指标限制,更贴合恢复完整相关提交集合的实际需求。我们对一系列涵盖主流技术路线的前沿模型进行了基准测试,包括基于规则、机器学习、深度学习以及基于Transformer的方法。实验结果表明,在平衡数据集上,现有模型仅能达到中等水平的性能,最高F1值(F1-score)仅为73%,这凸显了针对一对多关联恢复任务定制更高效模型的必要性。我们公开发布了本数据集与评估框架,以支持未来针对一对多关联恢复任务构建更鲁棒、更高效模型的相关研究。

提供机构:
Zenodo
创建时间:
2025-05-27
二维码
社区交流群
二维码
科研交流群
商业服务