An Empirical Study for Software Fault-Proneness Prediction with Ensemble Learning Models on Imbalanced Data Sets

An Empirical Study for Software Fault-Proneness Prediction with Ensemble Learning Models on Imbalanced Data Sets
复制标题

不平衡数据集上集成学习模型软件故障倾向预测的实证研究

DOI:
10.4304/jsw.9.3.697-704
复制
发表时间:
2014-01
期刊:
Journal of Software
影响因子:
--
通讯作者:
Renqing Li, Shihai Wang
Renqing Li, Shihai Wang
中科院分区:
其他
文献类型:
--
作者:
Renqing Li, Shihai Wang

文献摘要

参考文献

相似文献

软件故障会导致严重的系统错误和故障,造成巨大的经济损失。但是目前的检测和验证技术还不能发现和消除所有的软件错误。软件测试是检查这些错误和提高软件可靠性的重要手段,但显然这是一项非常昂贵的工作。模块的故障倾向性估计对于指导高风险模块的资源分配,从而最大限度地减少软件测试所需的资源是非常重要的。从而提高了软件测试的效率和软件的可靠性。然而,软件故障数据集本来就具有不均衡分布的特点。少量的软件模块包含大多数的故障,而大多数模块是无故障的。这种不平衡的数据分布对软件易错性预测领域的研究人员来说是一个挑战。本文基于软件度量,对C4. 5、SVM、KNN、Logistic、NaiveBayes、AdaBoost和SMOTEBoost等软件故障预测模型进行了研究。我们进行了实证研究,这些模型的有效性不平衡的软件故障数据集从美国宇航局的MDP。通过对实验结果的综合比较,发现SMOTEBoost模型在预测高风险软件模块方面表现突出,具有更高的查全率和AUC值,表明基于SMOTEBoost的模型具有更好的预测模块故障倾向性的能力,从而提高了软件测试的效率。
Software faults could cause serious system errors and failures, leading to huge economic losses. But currently none of inspection and verification technique is able to find and eliminate all software faults. Software testing is an important way to inspect these faults and raise software reliability, but obviously it is a really expensive job. The estimation of a module’s fault-proneness is important to minimize the software testing resources required by guiding the resource allocation on the high-risk modules. Consequently the efficiency of software testing and the reliability of the software are improved. The software faults data sets, however, originally have the imbalanced distribution. A small amount of software modules holds most faults, while the most of modules are fault-free. Such imbalanced data distribution is really a challenge for the researchers in the field of prediction for software fault-proneness. In this paper, we make an investigation on software fault-prone prediction models by employing C4.5, SVM, KNN, Logistic, NaiveBayes, AdaBoost and SMOTEBoost based on software metrics. We perform an empirical study on the effectiveness of these models on imbalanced software fault data sets obtained from NASA’s MDP. After a comprehensive comparison based on the experiment results, the SMOTEBoost reveals the outstanding performances than the other models on predicting the high-risk software modules with higher recall and AUC values, which demonstrates the model based on SMOTEBoost has a better ability to estimate a module’s fault-proneness and furthermore improve the efficiency of software testing.
DOI: --
发表时间: 2014
期刊: Journal of management science
影响因子: --
作者:
อนิรุธ สืบสิงห์
通讯作者: อนิรุธ สืบสิงห์
DOI: 10.1109/tse.2007.256941
发表时间: 2007
影响因子: 7.4
作者:
T. Menzies;Jeremy Greenwald;A. Frank
通讯作者: T. Menzies;Jeremy Greenwald;A. Frank
DOI: --
发表时间: 1992-10
期刊: --
影响因子: --
作者:
J. R. Quinlan
通讯作者: J. R. Quinlan
DOI: 10.1016/j.jss.2007.05.035
发表时间: 2008-02-01
影响因子: 3.5
作者:
Gondra, Iker
通讯作者: Gondra, Iker
DOI: 10.5121/ijcsa.2012.2203
发表时间: 2012-04
影响因子: --
作者:
C. Akalya;K. E. Kannammal;B Surendiran
通讯作者: C. Akalya;K. E. Kannammal;B Surendiran