A Uniform Architecture Design for Accelerating 2D and 3D CNNs on FPGAs

A Uniform Architecture Design for Accelerating 2D and 3D CNNs on FPGAs
复制标题

用于在 FPGA 上加速 2D 和 3D CNN 的统一架构设计

DOI:
10.3390/electronics8010065
复制
发表时间:
2019-01-01
期刊:
影响因子:
2.9
通讯作者:
Zhou, Jie
Zhou, Jie
中科院分区:
工程技术3区
文献类型:
--
作者:
Liu, Zhiqiang;Chow, Paul;Zhou, Jie

文献摘要

被引文献

相似文献

三维卷积神经网络(3D CNN)在许多复杂的计算机视觉应用中越来越受欢迎。许多基于FPGA的定制加速器被提出用于2D CNN,而很少用于3D CNN。三维CNN的计算量要大得多,并且由于引入了多一个维度,3D CNN加速的设计空间进一步扩大,这使得在FPGA上加速3D CNN成为一个巨大的挑战。由于发现2D和3D CNN的计算模式非常相似,我们在本文中提出了一种用于加速2D和3D CNN的统一架构设计。统一的架构是基于映射卷积矩阵乘法的想法。该算法采用自定义映射模块生成特征矩阵切片,无需将整个放大后的特征矩阵存储在片内或片外;采用分裂策略重构卷积层以适应片内存储容量;采用二维乘加(MAC)阵列高效计算矩阵乘法。为了演示,我们在Xilinx VC 709板上实现了一个具有高级合成(HLS)方法的加速器原型,并在三个典型的CNN模型上测试了加速器:AlexNet,VGG 16和C3 D。实验结果表明,该加速器在2D和3D CNN上都实现了最先进的吞吐量性能,其能效比CPU和GPU好得多。
Three-dimensional convolutional neural networks (3D CNNs) have gained popularity in many complicated computer vision applications. Many customized accelerators based on FPGAs are proposed for 2D CNNs, while very few are for 3D CNNs. Three-D CNNs are far more computationally intensive and the design space for 3D CNN acceleration has been further expanded since one more dimension is introduced, making it a big challenge to accelerate 3D CNNs on FPGAs. Motivated by the finding that the computation patterns of 2D and 3D CNNs are very similar, we propose a uniform architecture design for accelerating both 2D and 3D CNNs in this paper. The uniform architecture is based on the idea of mapping convolutions to matrix multiplications. A customized mapping module is developed to generate the feature matrix tilings with no need to store the entire enlarged feature matrix on-chip or off-chip, a splitting strategy is adopted to reconstruct a convolutional layer to adapt to the on-chip memory capacity, and a 2D multiply-and-accumulate (MAC) array is adopted to compute matrix multiplications efficiently. For demonstration, we implement an accelerator prototype with a high-level synthesis (HLS) methodology on a Xilinx VC709 board and test the accelerator on three typical CNN models: AlexNet, VGG16, and C3D. Experimental results show that the accelerator achieves state-of-the-art throughput performance on both 2D and 3D CNNs, with much better energy efficiency than the CPU and GPU.