Automated Program Repair in the Era of Large Pre-trained Language Models

Automated Program Repair in the Era of Large Pre-trained Language Models
复制标题

DOI:
10.1109/icse48619.2023.00129
复制
发表时间:
2023-05
期刊:
2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Chun Xia;Yuxiang Wei-;Lingming Zhang
Chun Xia;Yuxiang Wei-;Lingming Zhang
中科院分区:
其他
文献类型:
--
作者:
Chun Xia;Yuxiang Wei-;Lingming Zhang

文献摘要

被引文献

相似文献

自动程序修复(APR)旨在帮助开发人员自动修补软件错误。然而,目前最先进的传统和基于学习的APR技术面临的问题是补丁种类有限,无法修复复杂的错误。这主要是由于依赖bug修复数据集来制作修复模板(传统)或直接预测潜在的补丁(基于学习)。使用数十亿个文本/代码标记训练的大型预训练语言模型(LLM)可能有助于避免这个问题。最近,研究人员直接利用LLM进行APR,而不依赖于任何错误修复数据集。与此同时,这些现有的工作要么没有包括最先进的LLM,要么没有在现实的数据集上进行评估。因此,现代LLM在重要的APR问题上的真正力量还有待揭示。在这项工作中,我们进行了第一次广泛的研究,直接应用LLM的APR。我们选择了9个最近的国家的最先进的LLM,包括生成和填充模型,从125 M到20 B的大小。我们设计了3种不同的修复设置来评估我们可以使用LLM生成补丁的不同方式:1)生成整个补丁函数,2)填充给定前缀和后缀的代码块3)输出单行修复。我们将这些修复设置下的LLM应用于3种不同语言的5个数据集,并在修复的错误数量,生成速度和编译率方面比较不同的LLM。我们还比较了LLM对最近的国家的最先进的APR工具。我们的研究表明,直接应用最先进的LLM已经可以在我们所有的数据集上大大超过所有现有的APR技术。在所研究的LLM中,APR存在缩放效应,其中较大的模型倾向于实现更好的性能。此外,我们第一次表明,后缀代码后的错误行(采用渗透式APR)是重要的,不仅产生更多的修复,但更多的补丁,更高的编译率。除了补丁生成,LLM认为正确的补丁比其他补丁更自然,甚至可以用于有效的补丁排名或补丁正确性检查。最后,我们表明,基于LLM的APR可以通过以下方式进一步大幅提升:1)增加样本量,2)合并固定模板信息。
Automated Program Repair (APR) aims to help developers automatically patch software bugs. However, current state-of-the-art traditional and learning-based APR techniques face the problem of limited patch variety, failing to fix complicated bugs. This is mainly due to the reliance on bug-fixing datasets to craft fix templates (traditional) or directly predict potential patches (learning-based). Large Pre-Trained Language Models (LLMs), trained using billions of text/code tokens, can potentially help avoid this issue. Very recently, researchers have directly leveraged LLMs for APR without relying on any bug-fixing datasets. Meanwhile, such existing work either failed to include state-of-the-art LLMs or was not evaluated on realistic datasets. Thus, the true power of modern LLMs on the important APR problem is yet to be revealed. In this work, we perform the first extensive study on directly applying LLMs for APR. We select 9 recent state-of-the-art LLMs, including both generative and infilling models, ranging from 125M to 20B in size. We designed 3 different repair settings to evaluate the different ways we can use LLMs to generate patches: 1) generate the entire patch function, 2) fill in a chunk of code given the prefix and suffix 3) output a single line fix. We apply the LLMs under these repair settings on 5 datasets across 3 different languages and compare different LLMs in the number of bugs fixed, generation speed and compilation rate. We also compare the LLMs against recent state-of-the-art APR tools. Our study demonstrates that directly applying state-of-the-art LLMs can already substantially outperform all existing APR techniques on all our datasets. Among the studied LLMs, the scaling effect exists for APR where larger models tend to achieve better performance. Also, we show for the first time that suffix code after the buggy line (adopted in infilling-style APR) is important in not only generating more fixes but more patches with higher compilation rate. Besides patch generation, the LLMs consider correct patches to be more natural than other ones, and can even be leveraged for effective patch ranking or patch correctness checking. Lastly, we show that LLM-based APR can be further substantially boosted via: 1) increasing the sample size, and 2) incorporating fix template information.