遇见数据集

A Systematic Review of Unsupervised Defect Prediction Dataset

收藏
Mendeley Data2020-02-17 更新2026-04-09 收录
官方服务:

资源简介:

This dataset is about a systematic review of unsupervised learning techniques for software defect prediction (our related paper: "A Systematic Review of Unsupervised Learning Techniques for Software Defect Prediction" in Information and Software Technology [accepted in Feb, 2020] ). We conducted this systematic literature review that identified 49 studies which satisfied our inclusion criteria containing 2456 individual experimental results. In order to compare prediction performance across these studies in a consistent way, we recomputed the confusion matrices and employed MCC as our main performance measure. From each paper we extracted: Title, Year, Journal/conference, 'Predatory' publisher? (Y | N), Count of results reported in paper, Count of inconsistent results reported in paper, Parameter tuning in SDP? (Yes | Default | ?) and SDP references(SDPRefs OrigResults | SDPRefs |SDPNoRefs | OnlyUnSDP). Then from within each paper, we extracted for each experimental result including: Prediction method name (e.g., DTJ48), Project name trained on (e.g., PC4), Project name tested on (e.g., PC4), Prediction type (within-project | cross-project), No. of input metrics (count | NA), Dataset family (e.g., NASA), Dateset fault rate (%), Was cross validation used? (Y | N | ?), Was error checking possible? (Y | N), Inconsistent results? (Y | N | ?), Error reason description (text), Learning type (Supervised | Unsupervised), Clustering method? (Y | N | NA), Machine learning family (e.g., Un-NN), Machine learning technique (e.g., KM), Prediction results (including TP, TN, FP, FN, etc.).

本数据集围绕软件缺陷预测(Software Defect Prediction, SDP)的无监督学习技术系统性综述展开,相关研究论文为发表于《Information and Software Technology》、2020年2月录用的"A Systematic Review of Unsupervised Learning Techniques for Software Defect Prediction"。本次系统性文献综述共筛选出49项符合纳入标准的研究,涵盖2456项独立实验结果。为以统一标准对比各研究的预测性能,我们重新计算了混淆矩阵,并以马氏相关系数(Matthews Correlation Coefficient, MCC)作为核心性能评估指标。我们从每篇入选文献中提取了以下信息:文献标题、发表年份、刊载期刊/会议、是否为掠夺性出版商(Y|N)、文献报告的实验结果总数、文献报告的不一致结果数量、SDP参数调优情况(是|默认|?)以及SDP参考文献类型(SDPRefs OrigResults | SDPRefs | SDPNoRefs | OnlyUnSDP)。针对每篇文献中的每项实验结果,我们进一步提取了以下内容:预测方法名称(如DTJ48)、训练所用项目名称(如PC4)、测试所用项目名称(如PC4)、预测类型(项目内|跨项目)、输入指标数量(计数|NA)、数据集族(如NASA)、数据集缺陷率(%)、是否使用交叉验证(Y|N|?)、是否可开展误差检查(Y|N)、是否存在不一致结果(Y|N|?)、误差原因描述(文本)、学习类型(监督|无监督)、是否为聚类方法(Y|N|NA)、机器学习家族(如Un-NN)、机器学习技术(如KM)以及预测结果(包含TP、TN、FP、FN等指标)。

创建时间:
2020-02-17
二维码
社区交流群
二维码
科研交流群
商业服务