Large Language Models for Test-Free Fault Localization

Large Language Models for Test-Free Fault Localization
复制标题

DOI:
10.1145/3597503.3623342
复制
发表时间:
2023-10
期刊:
2024 IEEE/ACM 46th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Aidan Z. H. Yang;Ruben Martins;Claire Le Goues;Vincent J. Hellendoorn
Aidan Z. H. Yang;Ruben Martins;Claire Le Goues;Vincent J. Hellendoorn
中科院分区:
其他
文献类型:
--
作者:
Aidan Z. H. Yang;Ruben Martins;Claire Le Goues;Vincent J. Hellendoorn

文献摘要

被引文献

相似文献

故障定位(FL)旨在自动定位有错误的代码行,这是许多手动和自动调试任务的关键第一步。以前的FL技术假设提供输入测试,并且通常需要大量的程序分析、程序插装或数据预处理。之前针对APR的深度学习工作很难从小数据集中学习,并且在现实世界的程序中产生的结果有限。受大语言模型(LLM)的代码,以适应新的任务的能力很少的例子的基础上,我们调查的LLM行级故障定位的适用性。具体来说,我们建议克服LLM的从左到右的性质,通过微调LLM学习的表示上的一小组双向适配器层来产生LLMAO,这是第一种基于语言模型的故障定位方法,可以在没有任何测试覆盖信息的情况下定位错误代码行。我们用3.5亿、60亿和160亿个参数对LLM进行微调,这些参数是在小的、手动策划的有缺陷的程序语料库上进行的,比如$Defects4\mathcal{J}$语料库。我们观察到,我们的技术在建立在较大的模型上时,在故障定位方面实现了更大的信心,错误定位性能与LLM大小一致。我们的实证评估表明,LLMAO将最先进的机器学习故障定位(MLFL)基线的Top-1结果提高了2.3%-54.4%,Top-5结果提高了14.4%-35.6%。LLMAO也是第一个使用语言模型架构训练的FL技术,可以检测到代码行级别的安全漏洞。
Fault Localization (FL) aims to automatically localize buggy lines of code, a key first step in many manual and automatic debugging tasks. Previous FL techniques assume the provision of input tests, and often require extensive program analysis, program instrumentation, or data preprocessing. Prior work on deep learning for APR struggles to learn from small datasets and produces limited results on real-world programs. Inspired by the ability of large language models (LLMs) of code to adapt to new tasks based on very few examples, we investigate the applicability of LLMs to line level fault localization. Specifically, we propose to overcome the left-to-right nature of LLMs by fine-tuning a small set of bidirectional adapter layers on top of the representations learned by LLMs to produce LLMAO, the first language model based fault localization approach that locates buggy lines of code without any test coverage information. We fine-tune LLMs with 350 million, 6 billion, and 16 billion parameters on small, manually curated corpora of buggy programs such as the $Defects4\mathcal{J}$ corpus. We observe that our technique achieves substantially more confidence in fault localization when built on the larger models, with bug localization performance scaling consistently with the LLM size. Our empirical evaluation shows that LLMAO improves the Top-1 results over the state-of-the-art machine learning fault localization (MLFL) baselines by 2.3%-54.4%, and Top-5 results by 14.4%-35.6%. LLMAO is also the first FL technique trained using a language model architecture that can detect security vulnerabilities down to the code line level.