Improving the Accuracy of Spectrum-based Fault Localization for Automated Program Repair

Improving the Accuracy of Spectrum-based Fault Localization for Automated Program Repair
复制标题

DOI:
10.1145/3387904.3389290
复制
发表时间:
2020-07
期刊:
2020 IEEE/ACM 28th International Conference on Program Comprehension (ICPC)
影响因子:
--
通讯作者:
Tetsushi Kuma;Yoshiki Higo;S. Matsumoto;S. Kusumoto
Tetsushi Kuma;Yoshiki Higo;S. Matsumoto;S. Kusumoto
中科院分区:
其他
文献类型:
--
作者:
Tetsushi Kuma;Yoshiki Higo;S. Matsumoto;S. Kusumoto

文献摘要

相似文献

测试用例的充分性对于基于频谱的故障定位至关重要(简而言之,SBFL)。如果给定的一组测试用例不够,则SBFL不起作用。在这种情况下,我们可以通过添加新的测试案例来提高SBFL的可靠性。但是,在自动化程序维修的背景下(简称APR),添加许多未考虑其属性的测试案例不合适。例如,对于最著名的APR工具GenProg的情况,所有与错误模块有关的测试用例均针对每个突变程序执行。测试用例的执行结果用于检查它们是否通过所有测试用例并推断给定错误的错误陈述。因此,在APR的背景下,添加必要的最小测试用例以提高SBFL的准确性很重要。在本文中,我们提出了三种从大量自动生成的测试用例中选择一些测试用例的策略。我们在错误数据集缺陷4J上进行了一个小实验,并确认SBFL的准确性提高了56.3%的目标错误,而在最佳策略的情况下,准确性降低了17.3%。我们还确认,执行时间的增加被抑制至中位数1.5秒。
The sufficiency of test cases is essential for spectrum-based fault localization (in short, SBFL). If a given set of test cases is not sufficient, SBFL does not work. In such a case, we can improve the reliability of SBFL by adding new test cases. However, adding many test cases without considering their properties is not appropriate in the context of automated program repair (in short, APR). For example, in the case of GenProg, which is the most famous APR tool, all the test cases related to the bug module are executed for each of the mutated programs. Execution results of test cases are used for checking whether they pass all the test cases and inferring faulty statements for a given bug. Thus, in the context of APR, it is important to add necessary minimum test cases to improve the accuracy of SBFL. In this paper, we propose three strategies for selecting some test cases from a large number of automatically-generated test cases. We conducted a small experiment on bug dataset Defect4J and confirmed that the accuracy of SBFL was improved for 56.3% of target bugs while the accuracy was decreased for 17.3% in the case of the best strategy. We also confirmed that the increase of the execution time was suppressed to 1.5 seconds at the median.