Reducing Features to Improve Code Change-Based Bug Prediction

Reducing Features to Improve Code Change-Based Bug Prediction
复制标题

DOI:
10.1109/tse.2012.43
复制
发表时间:
2013-04
影响因子:
7.4
通讯作者:
S. Shivaji;James Whitehead;R. Akella;Sunghun Kim
S. Shivaji;James Whitehead;R. Akella;Sunghun Kim
中科院分区:
计算机科学1区
文献类型:
--
作者:
S. Shivaji;James Whitehead;R. Akella;Sunghun Kim

文献摘要

被引文献

相似文献

机器学习分类器最近作为一种方法出现,用于预测源代码文件更改中引入的错误。分类器首先根据软件历史进行训练,然后用于预测即将发生的更改是否会导致错误。现有的基于分类器的缺陷预测技术的缺点是性能不能满足实际应用,并且由于大量的机器学习特征而导致预测时间较慢。本文研究了通常适用于基于分类的缺陷预测方法的多种特征选择技术。在达到最佳分类性能之前,这些技术会丢弃不太重要的特征。用于训练的特征总数大幅减少,通常不到原始特征的10%。在11个软件项目上对朴素贝叶斯和支持向量机分类器的性能进行了表征。使用特征选择的朴素贝叶斯在错误F度量(改进21%)方面比先前的更改分类错误预测结果(由第二和第四作者[28])有显著改进。支持向量机在Buggy F度量上的改进为9%。有趣的是,对不同数量的功能的性能分析表明,即使只有原始功能数量的1%,也能实现强大的性能。
Machine learning classifiers have recently emerged as a way to predict the introduction of bugs in changes made to source code files. The classifier is first trained on software history, and then used to predict if an impending change causes a bug. Drawbacks of existing classifier-based bug prediction techniques are insufficient performance for practical use and slow prediction times due to a large number of machine learned features. This paper investigates multiple feature selection techniques that are generally applicable to classification-based bug prediction methods. The techniques discard less important features until optimal classification performance is reached. The total number of features used for training is substantially reduced, often to less than 10 percent of the original. The performance of Naive Bayes and Support Vector Machine (SVM) classifiers when using this technique is characterized on 11 software projects. Naive Bayes using feature selection provides significant improvement in buggy F-measure (21 percent improvement) over prior change classification bug prediction results (by the second and fourth authors [28]). The SVM's improvement in buggy F-measure is 9 percent. Interestingly, an analysis of performance for varying numbers of features shows that strong performance is achieved at even 1 percent of the original number of features.