Accelerating RTL simulation with GPUs

Accelerating RTL simulation with GPUs
复制标题

使用 GPU 加速 RTL 仿真

DOI:
10.1109/iccad.2011.6105404
复制
发表时间:
2011
期刊:
2011 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)
影响因子:
--
通讯作者:
Yangdong Deng
Yangdong Deng
中科院分区:
--
文献类型:
--
作者:
Hao Qian;Yangdong Deng

文献摘要

被引文献

相似文献

随着集成电路的快速复杂性,验证已成为当今IC设计流的瓶颈。实际上,在典型的IC设计项目中,可以将超过70%的IC设计转环时间用于验证过程。在各种验证任务中,寄存器传输级别(RTL)模拟是验证数字IC设计正确性的最广泛使用的方法。在模拟具有复杂内部行为的大型IC设计(例如,运行嵌入式软件的CPU内核)时,RTL模拟可能非常耗时。由于RTL-to-layout仍然是最普遍的IC设计方法,因此加速RTL仿真过程至关重要。最近,图形处理单元(GPGPU)上的通用计算正在成为加速计算密集型工作负载的有希望的范式。最近的一些作品证明了使用GPU加快门和系统级仿真任务的有效性。在这项工作中,我们提出了一个有效的GPU加速RTL模拟框架。我们引入了一种将Verilog RTL描述转化为等效GPU源代码的方法,以模拟GPU上的电路行为。另外,还采用了基于CMB的并行模拟协议来提供足够的并行性水平。由于RTL仿真缺乏数据级并行性,因此我们还提出了一种新的解决方案,可以将GPU用作有效的任务级并行处理器。实验结果证明,我们的基于GPU的模拟器的表现优于商业顺序RTL模拟器超过20倍。
With the fast increasing complexity of integrated circuits, verification has become the bottleneck of today's IC design flow. In fact, over 70% of the IC design turn-around time can be spent on the verification process in a typical IC design project. Among various verification tasks, Register Transfer Level (RTL) simulation is the most widely used method to validate the correctness of digital IC designs. When simulating a large IC design with complicated internal behaviors (e.g., CPU cores running embedded software), RTL simulation can be extremely time consuming. Since RTL-to-layout is still the most prevalent IC design methodology, it is essential to speedup the RTL simulation process. Recently, General Purpose computing on Graphics Processing Units (GPGPU) is becoming a promising paradigm to accelerate computing-intensive workloads. A few recent works have demonstrated the effectiveness of using GPU to expedite gate and system level simulation tasks. In this work, we proposed an efficient GPU-accelerated RTL simulation framework. We introduce a methodology to translate Verilog RTL description into equivalent GPU source code so as to simulate circuit behavior on GPUs. In addition, a CMB based parallel simulation protocol is also adopted to provide a sufficient level of parallelism. Because RTL simulation lacks data-level parallelism, we also present a novel solution to use GPU as an efficient task-level parallel processor. Experimental results prove that our GPU based simulator outperforms a commercial sequential RTL simulator by over 20 fold.