Evaluating and Improving Fault Localization

Evaluating and Improving Fault Localization
复制标题

DOI:
10.1109/icse.2017.62
复制
发表时间:
2017-05
期刊:
2017 IEEE/ACM 39th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Spencer Pearson;José Campos;René Just;G. Fraser;Rui Abreu;Michael D. Ernst;D. Pang;Benjamin Keller-Benjamin-K
Spencer Pearson;José Campos;René Just;G. Fraser;Rui Abreu;Michael D. Ernst;D. Pang;Benjamin Keller-Benjamin-K
中科院分区:
其他
文献类型:
--
作者:
Spencer Pearson;José Campos;René Just;G. Fraser;Rui Abreu;Michael D. Ernst;D. Pang;Benjamin Keller-Benjamin-K

文献摘要

被引文献

相似文献

大多数故障本地化技术都将输入作为一个错误的程序,并作为输出作为输出的可疑代码位置的排名列表,该程序可能有缺陷。当研究人员提出一种新的故障定位技术时,他们通常会在具有已知故障的程序上对其进行评估。该技术是根据其输出列表中出现有缺陷代码的位置对技术进行评分的。这样可以比较多个故障定位技术,以确定哪种更好。先前的研究已经使用人工断层评估了故障定位技术,该技术是由突变工具或手动产生的。换句话说,以前的研究确定了哪种故障定位技术最好在寻找人造故障方面。但是,尚不清楚哪种故障定位技术最好在寻找真正的故障方面。鉴于先前的工作表明人造故障与实际故障的相似之处和差异并不明显,答案是相同的。我们进行了一项复制研究,以评估文献中的10项主张,这些主张比较了故障定位技术(来自基于频谱和基于突变的家族)。我们在6个现实世界中使用了2995个人工故障。我们的结果支持以前的7个索赔中的7个具有统计学意义,但只有3个具有不可忽略的效应大小。然后,我们使用6个程序中的310个实际故障评估了相同的10个索赔。每个先前的结果都被驳斥或在统计和实践上都微不足道。我们的实验表明,人造故障对于预测哪些故障定位技术在实际故障上的表现不佳。鉴于这些结果,我们确定了一个设计空间,其中包括许多以前研究过的故障定位技术以及数百种新技术。我们使用一组395个实际故障的集合,通过实验确定设计空间中哪些因素最重要。然后,我们通过新技术扩展了这个设计空间。我们的几种新颖技术的表现优于所有现有技术,尤其是在排名前5个或前10名报告中对有缺陷的代码进行排名。
Most fault localization techniques take as input a faulty program, and produce as output a ranked list of suspicious code locations at which the program may be defective. When researchers propose a new fault localization technique, they typically evaluate it on programs with known faults. The technique is scored based on where in its output list the defective code appears. This enables the comparison of multiple fault localization techniques to determine which one is better. Previous research has evaluated fault localization techniques using artificial faults, generated either by mutation tools or manually. In other words, previous research has determined which fault localization techniques are best at finding artificial faults. However, it is not known which fault localization techniques are best at finding real faults. It is not obvious that the answer is the same, given previous work showing that artificial faults have both similarities to and differences from real faults. We performed a replication study to evaluate 10 claims in the literature that compared fault localization techniques (from the spectrum-based and mutation-based families). We used 2995 artificial faults in 6 real-world programs. Our results support 7 of the previous claims as statistically significant, but only 3 as having non-negligible effect sizes. Then, we evaluated the same 10 claims, using 310 real faults from the 6 programs. Every previous result was refuted or was statistically and practically insignificant. Our experiments show that artificial faults are not useful for predicting which fault localization techniques perform best on real faults. In light of these results, we identified a design space that includes many previously-studied fault localization techniques as well as hundreds of new techniques. We experimentally determined which factors in the design space are most important, using an overall set of 395 real faults. Then, we extended this design space with new techniques. Several of our novel techniques outperform all existing techniques, notably in terms of ranking defective code in the top-5 or top-10 reports.