How, and why, process metrics are better

How, and why, process metrics are better
复制标题

DOI:
10.1109/icse.2013.6606589
复制
发表时间:
2013-05
期刊:
2013 35th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Foyzur Rahman;Premkumar T. Devanbu
Foyzur Rahman;Premkumar T. Devanbu
中科院分区:
其他
文献类型:
--
作者:
Foyzur Rahman;Premkumar T. Devanbu

文献摘要

被引文献

相似文献

缺陷预测技术可以潜在地帮助我们将质量保证工作集中在最容易出现缺陷的文件上。现代统计工具使快速构建和部署预测模型变得非常容易。软件度量是预测模型的核心;了解不同类型的度量如何有效,尤其是为什么有效,对于成功的模型部署非常重要。在本文中,我们从几个不同的角度分析了过程和代码度量的适用性和有效性。我们在12个大型开源项目的85个版本中构建了许多预测模型,以解决不同指标集的性能、稳定性、可移植性和停滞不前问题。我们的结果表明,尽管在缺陷预测文献中广泛使用代码度量,但对于预测来说,代码度量通常不如过程度量有用。其次,我们发现代码度量具有很高的停滞性;它们在不同版本之间没有太大变化。这导致预测模型停滞不前,导致相同的文件被重复预测为缺陷;不幸的是,这些重复出现缺陷的文件最终证明缺陷密度相对较低。
Defect prediction techniques could potentially help us to focus quality-assurance efforts on the most defect-prone files. Modern statistical tools make it very easy to quickly build and deploy prediction models. Software metrics are at the heart of prediction models; understanding how and especially why different types of metrics are effective is very important for successful model deployment. In this paper we analyze the applicability and efficacy of process and code metrics from several different perspectives. We build many prediction models across 85 releases of 12 large open source projects to address the performance, stability, portability and stasis of different sets of metrics. Our results suggest that code metrics, despite widespread use in the defect prediction literature, are generally less useful than process metrics for prediction. Second, we find that code metrics have high stasis; they don't change very much from release to release. This leads to stagnation in the prediction models, leading to the same files being repeatedly predicted as defective; unfortunately, these recurringly defective files turn out to be comparatively less defect-dense.