遇见数据集

Replication package for Ticket-Augmented Just-in-Time Defect Prediction

收藏
Zenodo2025-06-17 更新2026-05-26 收录
官方服务:

资源简介:

The primary goal of bug prediction is to optimize testing efforts by focusing on software fragments, i.e., classes, methods, commits (JIT), or lines of code, most likely to be buggy. However, these predicted fragments already contain bugs. Thus, the current bug prediction approaches support fixing rather than prevention. The aim of this paper is to introduce and evaluate Ticket-Level Prediction (TLP), an approach to identify tickets that will introduce bugs once implemented. We analyze TLP at three temporal points, each point represents a ticket lifecycle stage: Open, In Progress, or Closed. We conjecture that: (1) TLP accuracy increases as tickets progress towards the closed stage due to improved feature reliability over time, and (2) the predictive power of features changes across these temporal points. Our TLP approach leverages 72 features belonging to six different families: code, developer, external temperature, internal temperature, intrinsic, ticket to tickets, and JIT. Our TLP evaluation uses a sliding-window approach, balancing feature selection and three machine-learning bug prediction classifiers on about 10,000 tickets of two Apache open-source projects. Our results show that TLP accuracy increases with proximity, confirming the expected trade-off between early prediction and accuracy. Regarding the prediction power of feature families, no single feature family dominates across stages; developer-centric signals are most informative early, whereas code and JIT metrics prevail near closure, and temperature-based features provide complementary value throughout. Our findings complement and extend the literature on bug prediction at the class, method, or commit level by showing that defect prediction can be effectively moved upstream, offering opportunities for risk-aware ticket triaging and developer assignment before any code is written.

缺陷预测(bug prediction)的核心目标,是通过聚焦于最易出现缺陷的软件片段——即类、方法、提交(JIT)或代码行——来优化测试工作投入。然而,这些被预测出的片段本身已存在缺陷,因此当前的缺陷预测方法仅支持缺陷修复,而非前置预防。本文的研究目标是介绍并评估工单级缺陷预测(Ticket-Level Prediction, TLP)方法,该方法可识别出在实施后将引入缺陷的工单。我们在三个时间节点上对TLP方法展开分析,每个节点对应工单生命周期的一个阶段:开放、进行中与已关闭。我们提出如下两项猜想:(1)随着工单逐步推进至已关闭阶段,由于特征可靠性随时间提升,TLP的预测准确率也会随之升高;(2)不同时间节点下,特征的预测能力会发生变化。我们的TLP方法共用到72项特征,涵盖6大类特征族:代码特征、开发者特征、外部热度特征、内部热度特征、固有特征、工单间关联特征以及JIT特征。TLP的评估采用滑动窗口方法,在两个Apache开源项目的约10000份工单数据集上,平衡了特征选择与三种机器学习缺陷预测分类器的实验设置。实验结果表明,TLP的准确率随时间节点愈发接近工单关闭阶段而提升,验证了早期预测与准确率之间存在预期中的权衡关系。就特征族的预测能力而言,并无某一类特征族在所有阶段均占据主导地位:早期以开发者相关特征的信息价值最高,临近关闭阶段则代码特征与JIT指标表现最优,而基于热度的特征则在全阶段均可提供互补价值。本研究的发现补充并拓展了现有针对类、方法或提交层面的缺陷预测相关研究:我们证明了缺陷预测可有效前移至代码编写前的阶段,为实现风险感知的工单分流与开发者分配提供了可行路径。

提供机构:
Zenodo
创建时间:
2025-06-17
二维码
社区交流群
二维码
科研交流群
商业服务