Benchmarking Large Language Models for Automated Verilog RTL Code Generation

Benchmarking Large Language Models for Automated Verilog RTL Code Generation
复制标题

DOI:
10.23919/date56975.2023.10137086
复制
发表时间:
2022-12
期刊:
2023 Design, Automation & Test in Europe Conference & Exhibition (DATE)
影响因子:
--
通讯作者:
Shailja Thakur;Baleegh Ahmad;Zhenxing Fan;H. Pearce;Benjamin Tan;R. Karri;Brendan Dolan-Gavitt;S. Garg
Shailja Thakur;Baleegh Ahmad;Zhenxing Fan;H. Pearce;Benjamin Tan;R. Karri;Brendan Dolan-Gavitt;S. Garg
中科院分区:
其他
文献类型:
--
作者:
Shailja Thakur;Baleegh Ahmad;Zhenxing Fan;H. Pearce;Benjamin Tan;R. Karri;Brendan Dolan-Gavitt;S. Garg

文献摘要

被引文献

相似文献

自动化硬件设计可以消除工程过程中的大量人为错误,从而减少错误。Verilog是一种流行的硬件描述语言,用于建模和设计数字系统,因此生成Verilog代码是关键的第一步。新兴的大型语言模型(LLM)能够用其他编程语言编写高质量的代码。在本文中,我们的LLM生成有用的Verilog的能力的特点。为此,我们在从GitHub和Verilog教科书收集的Verilog数据集上微调了预训练的LLM。我们构建了一个评估框架,包括功能分析测试台和测试Verilog代码的语法生成的不同难度的问题的流程。我们的研究结果表明,在我们的问题场景中,微调导致LLM更能够生成语法正确的代码(总体为25.9%)。此外,在分析功能正确性时,经过微调的开源CodeGen LLM可以超越最先进的商业Codex LLM(总体为6.5%)。我们将我们的培训/评估脚本和LLM检查点作为开源贡献发布。
Automating hardware design could obviate a signif-icant amount of human error from the engineering process and lead to fewer errors. Verilog is a popular hardware description language to model and design digital systems, thus generating Verilog code is a critical first step. Emerging large language models (LLMs) are able to write high-quality code in other programming languages. In this paper, we characterize the ability of LLMs to generate useful Verilog. For this, we fine-tune pre-trained LLMs on Verilog datasets collected from GitHub and Verilog textbooks. We construct an evaluation framework comprising test-benches for functional analysis and a flow to test the syntax of Verilog code generated in response to problems of varying difficulty. Our findings show that across our problem scenarios, the fine-tuning results in LLMs more capable of producing syntactically correct code (25.9% overall). Further, when analyzing functional correctness, a fine-tuned open-source CodeGen LLM can outperform the state-of-the-art commercial Codex LLM (6.5% overall). We release our training/evaluation scripts and LLM checkpoints as open source contributions.