遇见数据集
官方服务:

资源简介:

The performance of the defect prediction model by using balanced and imbalanced datasets makes a big impact on the discovery of future defects. Current resampling techniques only address the imbalanced datasets without taking into consideration redundancy and noise inherent to the imbalanced datasets. To address the imbalance issue, we propose Kernel Crossover Oversampling (KCO), an oversampling technique based on kernel analysis and crossover interpolation. Specifically, the proposed technique aims to generate balanced datasets by increasing data diversity in order to reduce redundancy and noise. KCO first represents multidimensional features into two-dimensional features by employing Kernel Principal Component Analysis (KPCA). KCO then divides the plotted data distribution by deploying spectral clustering to select the best region for interpolation. Lastly, KCO generates the new defect data by interpolating different data templates within the selected data clusters. According to the prediction evaluation conducted, KCO consistently produced F-scores ranging from 21% to 63% across six datasets, on average. According to the experimental results presented in this study, KCO provides more effective prediction performance than other baseline techniques. The experimental results show that KCO within project and cross project predictions especially consistently achieve higher performance of F-score results.

采用平衡与非平衡数据集训练的缺陷预测模型性能,对未来缺陷的发掘工作具有显著影响。现有重采样技术仅针对非平衡数据集开展处理,却未考虑此类数据集固有的冗余与噪声问题。为解决非平衡问题,本文提出核交叉过采样(Kernel Crossover Oversampling, KCO),一种基于核分析与交叉插值的过采样技术。具体而言,所提方法通过提升数据多样性生成平衡数据集,以降低数据冗余与噪声干扰。KCO首先借助核主成分分析(Kernel Principal Component Analysis, KPCA)将高维特征映射至二维特征空间;随后通过谱聚类划分已绘制的数据分布,选取最优插值区域;最后在选定的数据簇内对不同数据模板进行插值操作,生成新的缺陷样本数据。经开展的预测评估验证,KCO在6个数据集上的F值(F-score)始终处于21%至63%的区间内,整体平均性能表现优异。本研究的实验结果表明,相较于其他基线技术,KCO可实现更出色的预测性能。实验结果进一步显示,KCO在项目内与跨项目缺陷预测任务中,始终能取得更高的F值性能表现。

创建时间:
2024-04-11
二维码
社区交流群
二维码
科研交流群
商业服务