Benchmarking Large Language Models for Automated Verilog RTL Code Generation
Benchmarking Large Language Models for Automated Verilog RTL Code Generation
复制标题
DOI:
10.23919/date56975.2023.10137086
复制
发表时间:
2022-12
期刊:
影响因子:
--
通讯作者:
Shailja Thakur;Baleegh Ahmad;Zhenxing Fan;H. Pearce;Benjamin Tan;R. Karri;Brendan Dolan-Gavitt;S. Garg
中科院分区:
文献类型:
--
作者:
Shailja Thakur;Baleegh Ahmad;Zhenxing Fan;H. Pearce;Benjamin Tan;R. Karri;Brendan Dolan-Gavitt;S. Garg
Automating hardware design could obviate a signif-icant amount of human error from the engineering process and lead to fewer errors. Verilog is a popular hardware description language to model and design digital systems, thus generating Verilog code is a critical first step. Emerging large language models (LLMs) are able to write high-quality code in other programming languages. In this paper, we characterize the ability of LLMs to generate useful Verilog. For this, we fine-tune pre-trained LLMs on Verilog datasets collected from GitHub and Verilog textbooks. We construct an evaluation framework comprising test-benches for functional analysis and a flow to test the syntax of Verilog code generated in response to problems of varying difficulty. Our findings show that across our problem scenarios, the fine-tuning results in LLMs more capable of producing syntactically correct code (25.9% overall). Further, when analyzing functional correctness, a fine-tuned open-source CodeGen LLM can outperform the state-of-the-art commercial Codex LLM (6.5% overall). We release our training/evaluation scripts and LLM checkpoints as open source contributions.