Is the cure worse than the disease? overfitting in automated program repair

Is the cure worse than the disease? overfitting in automated program repair
复制标题

DOI:
10.1145/2786805.2786825
复制
发表时间:
2015-08
期刊:
Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering
影响因子:
--
通讯作者:
Edward K. Smith;Earl T. Barr;Claire Le Goues;Yuriy Brun
Edward K. Smith;Earl T. Barr;Claire Le Goues;Yuriy Brun
中科院分区:
其他
文献类型:
--
作者:
Edward K. Smith;Earl T. Barr;Claire Le Goues;Yuriy Brun

文献摘要

被引文献

相似文献

自动化程序维修已经显示出减少大量手动工作调试所需的有望。本文解决了通过维修程序引起的自动修复技术的早期评估的不足,并使用相同的测试评估了生成的补丁的正确性。由于测试是程序正确性的不完善度量,因此对此类评估不会区分正确的补丁和过度拟合可用测试并破坏未经测试但期望的功能的补丁。本文在公开可用的错误基准上评估了两种经过良好研究的维修工具,即GenProg和Trpautorepair,每个基准都带有人编写的补丁。通过使用独立于维修期间使用的测试评估贴片,我们发现工具不太可能改善通过的独立测试的比例,并且斑块的质量与维修过程中使用的测试套件的覆盖率成正比。对于通过大多数测试的程序,这些工具可能会破坏测试与修复测试一样。但是,新手开发人员也过度合适,自动化维修的性能并不比这些开发人员差。除了过度拟合之外,我们还测量了测试套件覆盖范围,测试套件出处和启动计划质量的影响,以及新手开发者写作和工具生成的补丁之间的质量差异。来自用于补丁生成的一个。
Automated program repair has shown promise for reducing the significant manual effort debugging requires. This paper addresses a deficit of earlier evaluations of automated repair techniques caused by repairing programs and evaluating generated patches' correctness using the same set of tests. Since tests are an imperfect metric of program correctness, evaluations of this type do not discriminate between correct patches and patches that overfit the available tests and break untested but desired functionality. This paper evaluates two well-studied repair tools, GenProg and TrpAutoRepair, on a publicly available benchmark of bugs, each with a human-written patch. By evaluating patches using tests independent from those used during repair, we find that the tools are unlikely to improve the proportion of independent tests passed, and that the quality of the patches is proportional to the coverage of the test suite used during repair. For programs that pass most tests, the tools are as likely to break tests as to fix them. However, novice developers also overfit, and automated repair performs no worse than these developers. In addition to overfitting, we measure the effects of test suite coverage, test suite provenance, and starting program quality, as well as the difference in quality between novice-developer-written and tool-generated patches when quality is assessed with a test suite independent from the one used for patch generation.