Cross-project defect prediction models: L'Union fait la force

Cross-project defect prediction models: L'Union fait la force
复制标题

DOI:
10.1109/csmr-wcre.2014.6747166
复制
发表时间:
2014-02
期刊:
2014 Software Evolution Week - IEEE Conference on Software Maintenance, Reengineering, and Reverse Engineering (CSMR-WCRE)
影响因子:
--
通讯作者:
Annibale Panichella;Rocco Oliveto;A. D. Lucia
Annibale Panichella;Rocco Oliveto;A. D. Lucia
中科院分区:
其他
文献类型:
--
作者:
Annibale Panichella;Rocco Oliveto;A. D. Lucia

文献摘要

被引文献

相似文献

现有的缺陷预测模型使用产品或过程度量和机器学习方法来识别易出现缺陷的源代码实体。不同的分类器(例如,线性回归、逻辑回归或分类树)在过去十年中已经被研究。迄今为止取得的成果有时是对比鲜明的,并没有显示出一个明确的赢家。在本文中,我们提出了一个实证研究,旨在统计分析不同的缺陷预测的等效性。我们还提出了一种组合方法,称为CODEP(COmbined DEfect Predictor),该方法采用不同机器学习技术提供的分类来提高易出现缺陷的实体的检测。这项研究是在10个开源软件系统和跨项目缺陷预测的背景下进行的,这是缺陷预测领域的主要挑战之一。结果的统计分析表明,所研究的分类器并不等价,它们可以相互补充。与独立的缺陷预测器相比,CODEP实现的上级预测精度也证实了这一点。
Existing defect prediction models use product or process metrics and machine learning methods to identify defect-prone source code entities. Different classifiers (e.g., linear regression, logistic regression, or classification trees) have been investigated in the last decade. The results achieved so far are sometimes contrasting and do not show a clear winner. In this paper we present an empirical study aiming at statistically analyzing the equivalence of different defect predictors. We also propose a combined approach, coined as CODEP (COmbined DEfect Predictor), that employs the classification provided by different machine learning techniques to improve the detection of defect-prone entities. The study was conducted on 10 open source software systems and in the context of cross-project defect prediction, that represents one of the main challenges in the defect prediction field. The statistical analysis of the results indicates that the investigated classifiers are not equivalent and they can complement each other. This is also confirmed by the superior prediction accuracy achieved by CODEP when compared to stand-alone defect predictors.