SequenceR: Sequence-to-Sequence Learning for End-to-End Program Repair

SequenceR: Sequence-to-Sequence Learning for End-to-End Program Repair
复制标题

DOI:
10.1109/tse.2019.2940179
复制
发表时间:
2021-09-01
影响因子:
7.4
通讯作者:
Monperrus, Martin
Monperrus, Martin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chen, Zimin;Kommrusch, Steve;Monperrus, Martin

文献摘要

被引文献

相似文献

本文介绍了一种基于序列到序列学习的端到端程序修复新方法。我们设计、实现并评估了一种名为 SequenceR 的技术,该技术基于源代码的序列到序列学习来修复错误。这种方法利用复制机制来克服大代码中出现的词汇量无限的问题。我们的系统是数据驱动的;我们在 35,578 个样本上对其进行了训练,这些样本是从开源软件库的提交中精心挑选出来的。我们在 4,711 个独立的真实错误修复以及用于程序修复研究的 Defects4J 基准上对 SequenceR 进行了评估。SequenceR 能够完美预测 950/4,711 个测试样本的修复线路,并为 Defects4J 基准中的 14 个错误找到正确的补丁。SequenceR 能够捕捉各种修复算子,而无需针对特定领域进行自上而下的设计。
This paper presents a novel end-to-end approach to program repair based on sequence-to-sequence learning. We devise, implement, and evaluate a technique, called SequenceR, for fixing bugs based on sequence-to-sequence learning on source code. This approach uses the copy mechanism to overcome the unlimited vocabulary problem that occurs with big code. Our system is data-driven; we train it on 35,578 samples, carefully curated from commits to open-source repositories. We evaluate SequenceR on 4,711 independent real bug fixes, as well on the Defects4J benchmark used in program repair research. SequenceR is able to perfectly predict the fixed line for 950/4,711 testing samples, and find correct patches for 14 bugs in Defects4J benchmark. SequenceR captures a wide range of repair operators without any domain-specific top-down design.