What is the Impact of Imbalance on Software Defect Prediction Performance?

What is the Impact of Imbalance on Software Defect Prediction Performance?
复制标题

DOI:
10.1145/2810146.2810150
复制
发表时间:
2015-10
期刊:
Proceedings of the 11th International Conference on Predictive Models and Data Analytics in Software Engineering
影响因子:
--
通讯作者:
Zaheed Mahmood;David Bowes;Peter Lane;T. Hall
Zaheed Mahmood;David Bowes;Peter Lane;T. Hall
中科院分区:
其他
文献类型:
--
作者:
Zaheed Mahmood;David Bowes;Peter Lane;T. Hall

文献摘要

被引文献

相似文献

软件缺陷预测性能在较大范围内变化。 Menzies表明有80%的回忆有天花板效应[8]。所使用的大多数数据集都高度不平衡。本文问,使用不同数据集(不同程度的失衡水平)对预测性能的经验效应是什么?我们使用以前对600个故障预测模型及其结果的荟萃分析合成的数据。将四个模型评估度量(Mathews相关系数(MCC),F-量度,精度和召回)与相应的数据不平衡比进行了比较。当数据不平衡时,软件缺陷预测研究的预测性能很低。随着数据变得越来越平衡,预测模型的预测性能从平均MCC的平均MCC提高,直到少数群体占数据集中20%的实例,MCC的平均值约为0.34。随着少数群体的比例增加超过20%,预测绩效不会显着增加。使用超过20%的有缺陷情况的数据集在使用MCC时没有对预测性能产生重大影响。我们得出的结论是,比较缺陷预测研究的结果应考虑到数据的不平衡。
Software defect prediction performance varies over a large range. Menzies suggested there is a ceiling effect of 80% Recall [8]. Most of the data sets used are highly imbalanced. This paper asks, what is the empirical effect of using different datasets with varying levels of imbalance on predictive performance? We use data synthesised by a previous meta-analysis of 600 fault prediction models and their results. Four model evaluation measures (the Mathews Correlation Coefficient (MCC), F-Measure, Precision and Recall) are compared to the corresponding data imbalance ratio. When the data are imbalanced, the predictive performance of software defect prediction studies is low. As the data become more balanced, the predictive performance of prediction models increases, from an average MCC of 0.15, until the minority class makes up 20% of the instances in the dataset, where the MCC reaches an average value of about 0.34. As the proportion of the minority class increases above 20%, the predictive performance does not significantly increase. Using datasets with more than 20% of the instances being defective has not had a significant impact on the predictive performance when using MCC. We conclude that comparing the results of defect prediction studies should take into account the imbalance of the data.