Large Language Models and Simple, Stupid Bugs

Large Language Models and Simple, Stupid Bugs
复制标题

DOI:
10.1109/msr59073.2023.00082
复制
发表时间:
2023-03
期刊:
2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR)
影响因子:
--
通讯作者:
Kevin Jesse;Toufique Ahmed;Prem Devanbu;Emily Morgan
Kevin Jesse;Toufique Ahmed;Prem Devanbu;Emily Morgan
中科院分区:
其他
文献类型:
--
作者:
Kevin Jesse;Toufique Ahmed;Prem Devanbu;Emily Morgan

文献摘要

相似文献

随着强大的神经语言模型的出现,基于人工智能的系统帮助开发人员完成编码任务正变得广泛可用;副驾驶就是这样一个系统。Copilot使用Codex,一种大型语言模型(LLM),以前面的“提示”为条件来完成代码。然而,Codex是在公共GitHub存储库上进行培训的,即可能包含错误和漏洞的代码。先前的研究[1],b[2]表明,Codex重现了培训中出现的漏洞。在这项研究中,我们检查Codex是如何容易产生一个有趣的错误类别,单语句错误,通常被称为简单,愚蠢的错误或sstub在MSR社区。我们发现Codex和类似的llm确实有助于避免一些sstub,但确实产生了已知的、一字不差的sstub,其可能性是已知的、一字不差的代码的两倍。我们探讨了法典生成的sstub的后果,并提出了避免策略,建议减少已知的逐字sstub的产生,并增加产生已知的逐字修复的可能性。
With the advent of powerful neural language models, AI-based systems to assist developers in coding tasks are becoming widely available; Copilot is one such system. Copilot uses Codex, a large language model (LLM), to complete code conditioned on a preceding "prompt". Codex, however, is trained on public GitHub repositories, viz., on code that may include bugs and vulnerabilities. Previous studies [1], [2] show Codex reproduces vulnerabilities seen in training. In this study, we examine how prone Codex is to generate an interesting bug category, single statement bugs, commonly referred to as simple, stupid bugs or SStuBs in the MSR community. We find that Codex and similar LLMs do help avoid some SStuBs, but do produce known, verbatim SStuBs as much as 2x as likely than known, verbatim correct code. We explore the consequences of the Codex generated SStuBs and propose avoidance strategies that suggest the possibility of reducing the production of known, verbatim SStubs, and increase the possibility of producing known, verbatim fixes.