遇见数据集

Replication package for Ticket-Augmented Just-in-Time Defect Prediction

收藏
Zenodo2025-06-17 更新2026-05-26 收录
官方服务:

资源简介:

The primary goal of bug prediction is to optimize testing efforts by focusing on software fragments, i.e., classes, methods, commits (JIT), or lines of code, most likely to be buggy. However, these predicted fragments already contain bugs. Thus, the current bug prediction approaches support fixing rather than prevention. The aim of this paper is to introduce and evaluate Ticket-Level Prediction (TLP), an approach to identify tickets that will introduce bugs once implemented. We analyze TLP at three temporal points, each point represents a ticket lifecycle stage: Open, In Progress, or Closed. We conjecture that: (1) TLP accuracy increases as tickets progress towards the closed stage due to improved feature reliability over time, and (2) the predictive power of features changes across these temporal points. Our TLP approach leverages 72 features belonging to six different families: code, developer, external temperature, internal temperature, intrinsic, ticket to tickets, and JIT. Our TLP evaluation uses a sliding-window approach, balancing feature selection and three machine-learning bug prediction classifiers on about 10,000 tickets of two Apache open-source projects. Our results show that TLP accuracy increases with proximity, confirming the expected trade-off between early prediction and accuracy. Regarding the prediction power of feature families, no single feature family dominates across stages; developer-centric signals are most informative early, whereas code and JIT metrics prevail near closure, and temperature-based features provide complementary value throughout. Our findings complement and extend the literature on bug prediction at the class, method, or commit level by showing that defect prediction can be effectively moved upstream, offering opportunities for risk-aware ticket triaging and developer assignment before any code is written.

缺陷预测(bug prediction)的核心目标,是通过聚焦于最有可能存在缺陷的软件片段——即类、方法、JIT提交(commits)或代码行——来优化测试工作量分配。然而,现有缺陷预测方法所定位的目标片段,往往已实际存在缺陷,因此这类方法更多用于辅助缺陷修复,而非前置预防。 本文旨在介绍并评估工单级缺陷预测(Ticket-Level Prediction, TLP)方法,该方法可识别出在实现阶段将引入缺陷的工单。我们在三个时间节点上对TLP进行分析,每个节点对应工单生命周期的一个阶段:待处理(Open)、进行中(In Progress)与已关闭(Closed)。 我们提出两项研究假设:(1)随着工单逐步推进至已关闭阶段,由于特征可靠性随时间提升,TLP的预测精度会随之上升;(2)不同时间节点下,特征的预测能力会发生变化。 我们的TLP方法共用到72个特征,涵盖六大类:代码特征、开发者特征、外部热度特征、内部热度特征、固有属性特征、工单间关联特征以及JIT特征。本次TLP评估采用滑动窗口实验范式,在两个Apache开源项目的约1万条工单数据集上,针对特征选择与三类机器学习缺陷预测分类器进行了均衡配置。 实验结果表明,TLP的预测精度随工单接近关闭阶段而提升,验证了早期预测与预测精度之间的既定权衡关系。关于特征家族的预测能力,没有单一特征家族在所有阶段占据主导地位:开发者相关特征在早期阶段提供的信息价值最高,而代码与JIT指标在接近关闭阶段时表现更优,基于热度的特征则在全流程中提供互补价值。 本研究的发现补充并拓展了现有针对类、方法或JIT提交层面的缺陷预测研究,证明缺陷预测可有效前移至代码编写前的阶段,为风险感知的工单分诊与开发者任务分配提供了前置机遇。

提供机构:
Zenodo
创建时间:
2025-06-17
二维码
社区交流群
二维码
科研交流群
商业服务