Researcher Bias: The Use of Machine Learning in Software Defect Prediction

Researcher Bias: The Use of Machine Learning in Software Defect Prediction
复制标题

DOI:
10.1109/tse.2014.2322358
复制
发表时间:
2014-06
影响因子:
7.4
通讯作者:
M. Shepperd;David Bowes;T. Hall
M. Shepperd;David Bowes;T. Hall
中科院分区:
计算机科学1区
文献类型:
--
作者:
M. Shepperd;David Bowes;T. Hall

文献摘要

被引文献

相似文献

背景。预测易缺陷软件组件的能力将是有价值的。因此,有许多经验研究来评估有效实现这一目标的不同技术的性能。但是,没有一种技术主导,因此设计可靠的缺陷预测模型仍然存在问题。客观的。我们试图理解许多相互矛盾的实验结果,并了解哪些因素对预测性能具有最大的影响。方法。我们对所有相关的高质量主要研究缺陷预测进行荟萃分析,以确定哪些因素会影响预测性能。这是基于满足我们的纳入标准的42项主要研究,该研究共同报告了600组经验预测结果。通过逆向工程一个通用响应变量,我们构建一个随机效应方差分析模型,以检查四个模型构建因子(分类器,数据集,输入指标和研究人员组)的相对贡献,以模拟预测性能。结果。令人惊讶的是,我们发现分类器的选择对性能几乎没有影响(1.3%),相反,主要的解释因素是研究人员组。与做什么相比,谁完成工作的人重要。结论。为了克服这种高水平的研究人员偏见,缺陷预测研究人员应(i)进行盲目分析,(ii)改善报告协议,(iii)进行更多的组间研究以减轻专业知识问题。最后,需要进行研究以确定这种偏见是否在其他应用领域中很普遍。
Background. The ability to predict defect-prone software components would be valuable. Consequently, there have been many empirical studies to evaluate the performance of different techniques endeavouring to accomplish this effectively. However no one technique dominates and so designing a reliable defect prediction model remains problematic. Objective. We seek to make sense of the many conflicting experimental results and understand which factors have the largest effect on predictive performance. Method. We conduct a meta-analysis of all relevant, high quality primary studies of defect prediction to determine what factors influence predictive performance. This is based on 42 primary studies that satisfy our inclusion criteria that collectively report 600 sets of empirical prediction results. By reverse engineering a common response variable we build a random effects ANOVA model to examine the relative contribution of four model building factors (classifier, data set, input metrics and researcher group) to model prediction performance. Results. Surprisingly we find that the choice of classifier has little impact upon performance (1.3 percent) and in contrast the major (31 percent) explanatory factor is the researcher group. It matters more who does the work than what is done. Conclusion. To overcome this high level of researcher bias, defect prediction researchers should (i) conduct blind analysis, (ii) improve reporting protocols and (iii) conduct more intergroup studies in order to alleviate expertise issues. Lastly, research is required to determine whether this bias is prevalent in other applications domains.