FPGA-Based High-Throughput CNN Hardware Accelerator With High Computing Resource Utilization Ratio

FPGA-Based High-Throughput CNN Hardware Accelerator With High Computing Resource Utilization Ratio
复制标题

基于 FPGA 的_高吞吐量_CNN_硬件_加速器_With_高_计算_资源_利用率_比率

DOI:
10.1109/tnnls.2021.3055814
复制
发表时间:
2021-02-12
影响因子:
10.4
通讯作者:
Huang, Yihua
Huang, Yihua
中科院分区:
计算机科学1区
文献类型:
--
作者:
Huang, Wenjin;Wu, Huangtao;Huang, Yihua

文献摘要

被引文献

相似文献

基于现场可编程门阵列(FPGA)的CNN硬件加速器采用单计算引擎(CE)架构或多CE架构,近年来引起了极大的关注。加速器的实际吞吐量也越来越高,但由于计算资源映射机制和数据供应问题等原因,其吞吐量仍远低于理论吞吐量,为此提出了一种新型的复合硬件CNN加速器架构。为了有效地执行卷积层(CL),提出了一种新的基于行级流水线流策略的多CE架构。针对每个CE,提出了一种优化的映射机制以提高其计算资源的利用率,并设计了一种高效的数据系统以避免CE的空闲状态,并提供连续的数据供应。此外,为了减轻片外带宽压力,提出了一种加权数据分配策略。为了实现全连接层(FCL),提出了一种基于批处理计算方法的单CE架构。基于这些设计方法和策略,在XC 7VX 980 T FPGA平台上实现了VGG-16和ResNet-101。VGG-16加速器消耗了3395个乘法器,在150 MHz下获得了1 TOPS的吞吐量,即理论吞吐量(2 × 3395 × 150 MOPS)的98.15%。类似地,ResNet-101加速器在100 MHz下实现了600 GOPS,约为理论吞吐量(2 x3121 x 100 MOPS)的96.12%。
The field-programmable gate array (FPGA)-based CNN hardware accelerator adopting single-computing-engine (CE) architecture or multi-CE architecture has attracted great attention in recent years. The actual throughput of the accelerator is also getting higher and higher but is still far below the theoretical throughput due to the inefficient computing resource mapping mechanism and data supply problem, and so on. To solve these problems, a novel composite hardware CNN accelerator architecture is proposed in this article. To perform the convolution layer (CL) efficiently, a novel multiCE architecture based on a row-level pipelined streaming strategy is proposed. For each CE, an optimized mapping mechanism is proposed to improve its computing resource utilization ratio and an efficient data system with continuous data supply is designed to avoid the idle state of the CE. Besides, to relieve the off-chip bandwidth stress, a weight data allocation strategy is proposed. To perform the fully connected layer (FCL), a single-CE architecture based on a batch-based computing method is proposed. Based on these design methods and strategies, visual geometry group network-16 (VGG-16) and ResNet-101 are both implemented on the XC7VX980T FPGA platform. The VGG-16 accelerator consumed 3395 multipliers and got the throughput of 1 TOPS at 150 MHz, that is, about 98.15% of the theoretical throughput (2 x 3395 x150 MOPS). Similarly, the ResNet-101 accelerator achieved 600 GOPS at 100 MHz, about 96.12% of the theoretical throughput (2 x3121 x 100 MOPS).