The impact of tangled code changes on defect prediction models

The impact of tangled code changes on defect prediction models
复制标题

DOI:
10.1007/s10664-015-9376-6
复制
发表时间:
2016-04-01
影响因子:
4.1
通讯作者:
Zeller, Andreas
Zeller, Andreas
中科院分区:
计算机科学2区
文献类型:
--
作者:
Herzig, Kim;Just, Sascha;Zeller, Andreas

文献摘要

被引文献

相似文献

在与源代码控制管理系统交互时,开发人员经常会在单个事务中提交不相关或松散关联的代码变更。在分析版本历史时,这种纠缠不清的变更会使所有模块的所有变更看起来都是相关的,可能会因噪声和偏差而影响分析结果。在对五个开源 Java 项目的调查中,我们发现 7% 到 20% 的错误修复都是由多个纠缠在一起的变更组成的。我们使用多预测因子方法来解开变更,结果表明平均至少有 16.6% 的源文件与错误报告关联错误。这些错误的错误文件关联似乎不会对将源文件分类为至少有一个错误或没有错误的模型产生重大影响。但我们的实验表明,与在纠缠不清的错误数据集上训练和测试的模型相比,解开纠缠不清的代码变更可以产生更准确的回归错误预测模型--在我们的实验中,统计意义上的准确率提高在 5 % 到 200 % 之间。我们建议更好地组织变更,以限制缠结变更的影响。
When interacting with source control management system, developers often commit unrelated or loosely related code changes in a single transaction. When analyzing version histories, such tangled changes will make all changes to all modules appear related, possibly compromising the resulting analyses through noise and bias. In an investigation of five open-source Java projects, we found between 7 % and 20 % of all bug fixes to consist of multiple tangled changes. Using a multi-predictor approach to untangle changes, we show that on average at least 16.6 % of all source files are incorrectly associated with bug reports. These incorrect bug file associations seem to not significantly impact models classifying source files to have at least one bug or no bugs. But our experiments show that untangling tangled code changes can result in more accurate regression bug prediction models when compared to models trained and tested on tangled bug datasets-in our experiments, the statistically significant accuracy improvements lies between 5 % and 200 %. We recommend better change organization to limit the impact of tangled changes.