FPGA-Based Training Accelerator Utilizing Sparseness of Convolutional Neural Network

FPGA-Based Training Accelerator Utilizing Sparseness of Convolutional Neural Network
复制标题

DOI:
10.1109/fpl.2019.00036
复制
发表时间:
2019-09
期刊:
2019 29th International Conference on Field Programmable Logic and Applications (FPL)
影响因子:
--
通讯作者:
Hiroki Nakahara;Youki Sada;Masayuki Shimoda;Kouki Sayama;Akira Jinguji;Shimpei Sato
Hiroki Nakahara;Youki Sada;Masayuki Shimoda;Kouki Sayama;Akira Jinguji;Shimpei Sato
中科院分区:
其他
文献类型:
--
作者:
Hiroki Nakahara;Youki Sada;Masayuki Shimoda;Kouki Sayama;Akira Jinguji;Shimpei Sato

文献摘要

相似文献

卷积神经网络(cnn)的训练几乎完全是在大型gpu集群上进行的。然而,它消耗大量的电力。因此,需要具有低功耗优势的高速训练系统。本文提出了一种基于fpga的训练加速器,利用CNN的稀疏性,由通用卷积单元和分布式堆栈池化单元组成。提出的通用卷积架构支持各种卷积操作,例如现代CNN中使用的点卷积、深度卷积、大核卷积和亚特鲁卷积。此外,我们使用了一种微调方案,该方案加载预训练的密集CNN以减少训练过程的内存大小,同时考虑重要的连通性以保持识别准确性。我们的训练方案减少了85%的参数,加速了训练计算,减小了片上尺寸。因此,它消除了耗能的DRAM访问。我们在Xilinx Virtex UltraScale+ VC1525加速开发板上实现了所提出的培训加速器。实验结果表明,与现有NVIDIA RTX2080Ti GPU相比,本文提出的稀疏CNN训练加速器在FPGA上的速度提高了4倍,功耗降低了2.9倍,单功耗性能提高了11.6倍。
Training of convolutional neural networks (CNNs) is almost exclusively performed on large clusters of GPUs. However, it consumes vast amounts of power. Thus, high-speed training systems superior in low-power consumption are desired. This paper proposes an FPGA-based training accelerator utilizing a sparseness of a CNN, which consists of universal convolutional units and pooling units with distributed stacks. The proposed universal convolution architecture supports various convolution operations, such as the point-wise, depth-wise, large kernel and atrous convolutions used in the modern CNN. Additionally, we utilize a fine-tuning scheme, which loads a pre-trained dense CNN to reduce the memory size for the training process, while it considers important connectivity to preserve recognition accuracy. Our training scheme reduces 85% parameters to accelerate the training computation and reduce on-chip size. Thus, it eliminates energy-consuming DRAM accesses. We implemented the proposed training accelerator on a Xilinx Virtex UltraScale+ VC1525 acceleration development board. Experimental results show that the proposed sparse CNN training accelerator on the FPGA can achieve four times faster, 2.9 times lower power consumption, and 11.6 times better performance per power, compared to the existing NVIDIA RTX2080Ti GPU.