Applying machine learning to software fault-proneness prediction

Applying machine learning to software fault-proneness prediction
复制标题

DOI:
10.1016/j.jss.2007.05.035
复制
发表时间:
2008-02-01
影响因子:
3.5
通讯作者:
Gondra, Iker
Gondra, Iker
中科院分区:
计算机科学2区
文献类型:
--
作者:
Gondra, Iker

文献摘要

被引文献

相似文献

软件测试对质量保证的重要性怎么强调都不过分。模块故障倾向性的估计对于最小化成本和提高软件测试过程的有效性非常重要。不幸的是,没有通用的技术来估计软件的故障倾向。一些软件度量和故障倾向之间的相关性导致了各种基于多个度量的预测模型。许多工作都集中在如何选择最有可能表明故障倾向的软件度量。在本文中,我们提出使用机器学习来实现这一目的。具体地,给定关于软件度量值和报告错误的数量的历史数据,训练人工神经网络(ANN)。然后,为了确定每个软件度量在预测故障倾向性的重要性,对训练的ANN进行灵敏度分析。被认为是最关键的软件度量,然后用作基于ANN的故障倾向性的连续测量的预测模型的基础。我们还将故障倾向预测视为二元分类任务(即,模块可以包含错误或没有错误),并使用支持向量机(SVM)作为最先进的分类方法。我们进行了一个比较实验研究的有效性的ANN和SVM的数据集从美国宇航局的CANDData程序数据存储库。(c)2007年爱思唯尔公司All rights reserved.
The importance of software testing to quality assurance cannot be overemphasized. The estimation of a module's fault-proneness is important for minimizing cost and improving the effectiveness of the software testing process. Unfortunately, no general technique for estimating software fault-proneness is available. The observed correlation between some software metrics and fault-proneness has resulted in a variety of predictive models based on multiple metrics. Much work has concentrated on how to select the software metrics that are most likely to indicate fault-proneness. In this paper, we propose the use of machine learning for this purpose. Specifically, given historical data on software metric values and number of reported errors, an Artificial Neural Network (ANN) is trained. Then, in order to determine the importance of each software metric in predicting fault-proneness, a sensitivity analysis is performed on the trained ANN. The software metrics that are deemed to be the most critical are then used as the basis of an ANN-based predictive model of a continuous measure of fault-proneness. We also view fault-proneness prediction as a binary classification task (i.e., a module can either contain errors or be error-free) and use Support Vector Machines (SVM) as a state-of-the-art classification method. We perform a comparative experimental study of the effectiveness of ANNs and SVMs on a data set obtained from NASA's Metrics Data Program data repository. (c) 2007 Elsevier Inc. All rights reserved.