MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation

MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation
复制标题

DOI:
10.1109/tse.2023.3267446
复制
发表时间:
2023-07
影响因子:
7.4
通讯作者:
Federico Cassano;John Gouwar;Daniel Nguyen;S. Nguyen;Luna Phipps-Costin;Donald Pinckney;Ming-Ho Yee-M
Federico Cassano;John Gouwar;Daniel Nguyen;S. Nguyen;Luna Phipps-Costin;Donald Pinckney;Ming-Ho Yee-M
中科院分区:
计算机科学1区
文献类型:
--
作者:
Federico Cassano;John Gouwar;Daniel Nguyen;S. Nguyen;Luna Phipps-Costin;Donald Pinckney;Ming-Ho Yee-M

文献摘要

被引文献

相似文献

大型语言模型已经证明了生成自然语言和编程语言文本的能力。虽然当代的代码生成模型是在几种编程语言的语料库上训练的,但它们是使用通常是单语的基准测试来测试的。最广泛使用的代码生成基准只针对Python,因此几乎没有量化的证据表明代码生成模型在其他编程语言上的表现。我们提出了MultiPL-E,一个将单元测试驱动的代码生成基准转换为新语言的系统。我们通过使用MultiPL-E将两个流行的Python代码生成基准转换为18种其他编程语言,创建了第一个大规模多语言代码生成基准。我们使用MultiPL-E来扩展HumanEval基准(Chen等人,2021)和MBPP基准(Austin等人,2021年)到18种语言,包括一系列编程范式和流行程度。使用这些新的并行基准测试,我们评估了三种最先进的代码生成模型的多语言性能:Codex(Chen等人,2021)、CodeGen(Nijkamp等人,2022)和InCoder(Fried等人,2022年)。我们发现Codex在其他几种语言的Python上的性能相当甚至超过了它。MultiPL-E中表示的编程语言范围使我们能够探索语言频率和语言特征对模型性能的影响。最后,将代码生成基准编译为新编程语言的MultiPL-E方法是可伸缩和可扩展的,使得评估新模型,基准和语言变得简单。
Large language models have demonstrated the ability to generate both natural language and programming language text. Although contemporary code generation models are trained on corpora with several programming languages, they are tested using benchmarks that are typically monolingual. The most widely used code generation benchmarks only target Python, so there is little quantitative evidence of how code generation models perform on other programming languages. We propose MultiPL-E, a system for translating unit test-driven code generation benchmarks to new languages. We create the first massively multilingual code generation benchmark by using MultiPL-E to translate two popular Python code generation benchmarks to 18 additional programming languages. We use MultiPL-E to extend the HumanEval benchmark (Chen et al., 2021) and MBPP benchmark (Austin et al., 2021) to 18 languages that encompass a range of programming paradigms and popularity. Using these new parallel benchmarks, we evaluate the multi-language performance of three state-of-the-art code generation models: Codex (Chen et al., 2021), CodeGen (Nijkamp et al., 2022) and InCoder (Fried et al., 2022). We find that Codex matches or even exceeds its performance on Python for several other languages. The range of programming languages represented in MultiPL-E allow us to explore the impact of language frequency and language features on model performance. Finally, the MultiPL-E approach of compiling code generation benchmarks to new programming languages is both scalable and extensible, making it straightforward to evaluate new models, benchmarks, and languages.