遇见数据集

Defect Prediction: Xerces

收藏
Zenodo2020-09-18 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Background: This paper describes an analysis that was conducted on newly collected repository with 92 versions of 38 proprietary, open-source and academic projects. A preliminary study performed before showed the need for a further in-depth analysis in order to identify project clusters. <br> Aims: The goal of this research is to perform clustering on software projects in order to identify groups of software projects with similar characteristic from the defect prediction point of view. One defect prediction model should work well for all projects that belong to such group. The existence of those groups was investigated with statistical tests and by comparing the mean value of prediction efficiency. <br> Method: Hierarchical and k-means clustering, as well as Kohonen’s neural network was used to find groups of similar projects. The obtained clusters were investigated with the discriminant analysis. For each of the identified group a statistical analysis has been conducted in order to distinguish whether this group really exists. Two defect prediction models were created for each of the identified groups. The first one was based on the projects that belong to a given group, and the second one - on all the projects. Then, both models were applied to all versions of projects from the investigated group. If the predictions from the model based on projects that belong to the identified group are significantly better than the all-projects model (the mean values were compared and statistical tests were used), we conclude that the group really exists. <br> Results: Six different clusters were identified and the existence of two of them was statistically proven: 1) cluster proprietary B – T=19, p=0.035, r=0.40; 2) cluster proprietary/open - t(17)=3.18, p=0.05, r=0.59. The obtained effect sizes (r) represent large effects according to Cohen’s benchmark, which is a substantial finding. <br> Conclusions: The two identified clusters were described and compared with results obtained by other researchers. The results of this work makes next step towards defining formal methods of reuse defect prediction models by identifying groups of projects within which the same defect prediction model may be used. Furthermore, a method of clustering was suggested and applied.

研究背景:本文针对新采集的软件仓库展开分析,该仓库涵盖38个专有、开源及学术项目的共计92个版本。此前开展的一项初步研究表明,为识别项目聚类,需开展进一步的深度分析。 研究目标:本研究旨在对软件项目进行聚类分析,以从缺陷预测视角识别出具有相似特征的软件项目群组。理想状态下,同一群组内的所有项目均可适配同一套缺陷预测模型。本研究通过统计检验及对比预测效率均值的方式,对这类群组的存在性展开了验证。 研究方法:本研究采用层次聚类、k-means聚类以及科赫农神经网络(Kohonen Neural Network)来挖掘相似项目群组。针对得到的聚类结果,采用判别分析进行验证。针对每一个识别出的群组,均开展了统计分析以确认该群组的真实存在性。针对每个识别出的群组,均构建了两套缺陷预测模型:第一套基于该群组内的项目进行训练,第二套则基于全部调研项目进行训练。随后,将两套模型分别应用于该群组内所有版本的项目。若基于目标群组项目训练的模型,其预测效果显著优于基于全项目训练的模型(通过对比均值并开展统计检验验证),则可判定该群组真实存在。 研究结果:本研究共识别出6个不同的聚类,其中2个聚类的存在性得到了统计验证:1)专有项目聚类B(T=19,p=0.035,r=0.40);2)专有/开源项目聚类(t(17)=3.18,p=0.05,r=0.59)。根据科恩(Cohen)效应量基准,本次得到的效应量(r)属于大效应量,这是一项极具价值的发现。 研究结论:本研究对识别出的两个聚类进行了描述,并与其他研究者的成果进行了对比。本研究成果通过识别出可复用同一套缺陷预测模型的项目群组,为定义缺陷预测模型的形式化复用方法迈出了关键一步。此外,本研究提出并应用了一套项目聚类方法。

提供机构:
Zenodo
创建时间:
2017-02-03
二维码
社区交流群
二维码
科研交流群
商业服务