APCNN: Explore Multi-Layer Cooperation for CNN Optimization and Acceleration on FPGA

APCNN: Explore Multi-Layer Cooperation for CNN Optimization and Acceleration on FPGA
复制标题

APCNN:探索多层合作在 FPGA 上实现 CNN 优化和加速

DOI:
10.1145/3431920.3439461
复制
发表时间:
2021
期刊:
The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays
影响因子:
--
通讯作者:
Fu, Song
Fu, Song
中科院分区:
--
文献类型:
--
作者:
Jiang, Beilei;Cheng, Xianwei;Tang, Sihai;Ma, Xu;Gu, Zhaochen;Zhao, Hui;Fu, Song

文献摘要

相似文献

本文介绍了APCNN,它探索了算法-硬件协同设计,并提供了一个基于FPGA的多层协同优化和定制设计的CNN加速框架。在算法设计方面,池化层被移动到APCNN中的非线性激活函数和归一化之前,我们证明了这会导致可以忽略不计的精度损失;然后通过冗余乘法消除,局部加法重用和全局加法重用来共同优化池化层与卷积层。我们进一步设计了一个专用的加速器,充分利用卷积池的跨层优化,不仅加快计算,但也减少了在FPGA上的芯片上的数据通信。我们证明,我们的新的APCNN可以实现75%的乘法和75%的加法减少在最好的情况下。对于通断芯片数据通信,可以消除存储器占用的最大百分比{Row,Col} /(Row × Col),其中Row和Col分别是激活特征图中的行数和列数。我们实现了APCNN的原型,并使用加速器级循环和能量模型以及RTL实现在LeNet-5和VGG 16上评估了其性能。我们的实验结果表明,与密集CNN相比,APCNN实现了2.5倍的加速比和4.7倍的能量效率。(This研究部分由NSF赠款CCF-1563750,OAC-2017564和CNS-2037982支持。
In this paper, we introduce APCNN, which explores algorithm-hardware co-design and provides a CNN acceleration framework with multi-layer cooperative optimization and customized design on FPGA. In terms of the algorithm design, the pooling layer is moved before the non-linear activation function and normalization in APCNN, which we prove causes negligible accuracy loss; the pooling layer is then co-optimized with the convolutional layer by means of redundant multiplication elimination, local addition reuse, and global addition reuse. We further design a dedicated accelerator to take full advantage of convolutional-pooling cross-layer optimization to not only accelerate computation but also reduce on-off chip data communication on FPGA. We demonstrate that our novel APCNN can achieve 75% multiplication and 75% addition reduction in the best case. For on-off chip data communication, a max{Row,Col} /(Row x Col) percent of memory footprint can be eliminated, where Row and Col are the number of rows and columns in the activation feature map respectively. We have implemented a prototype of APCNN and evaluated its performance on LeNet-5 and VGG16 using both an accelerator-level cycle and energy model and an RTL implementation. Our experimental results show that APCNN achieves a 2.5× speedup and 4.7× energy efficiency compared with the dense CNN. (This research was supported in part by NSF grants CCF-1563750, OAC-2017564, and CNS-2037982.)