Power-Efficient Deep Convolutional Neural Network Design Through Zero-Gating PEs and Partial-Sum Reuse Centric Dataflow

Power-Efficient Deep Convolutional Neural Network Design Through Zero-Gating PEs and Partial-Sum Reuse Centric Dataflow
复制标题

DOI:
10.1109/access.2021.3053259
复制
发表时间:
2021
期刊:
影响因子:
3.9
通讯作者:
Lin Ye;Jinghao Ye;M. Yanagisawa;Youhua Shi
Lin Ye;Jinghao Ye;M. Yanagisawa;Youhua Shi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Lin Ye;Jinghao Ye;M. Yanagisawa;Youhua Shi

文献摘要

相似文献

卷积神经网络(CNN)在目标检测和模式识别等许多领域取得了巨大的成功,但其代价是极高的计算复杂度和大量的外部存储器访问,这使得最先进的深度CNN难以在电池容量有限的资源受限的便携式/可穿戴设备上实现。为了解决这一设计挑战,本文提出了一种通过零门控处理元件(PE)和以部分和重用为中心的低功耗CNN设计。不同于现有的作品,要么只考虑激活地图中的零或使用片外训练过程的片上计算减少,提出了一种零门控PE设计,以避免不必要的片上计算,利用大量的零在两个过滤器的权重的预训练模型和激活地图。此外,还提出了一种以部分和重用为中心的DRAM访问减少算法。评估结果表明,与基线PE设计和现有的仅激活门控设计(即在Eyeriss)相比,我们的建议的PE阵列的整体功耗可以分别降低37%和14%,在8%和1%的面积开销的成本。此外,所提出的方法可以实现35%和47%的DRAM访问减少与相应的14%和49%的能源节省AlexNet和VGG-16相比,在Eyeriss。
Convolution neural networks (CNNs) have shown great success in many areas such as object detection and pattern recognition at the cost of extreme high computation complexity and significant external memory access, which makes state-of-the-art deep CNNs difficult to be implemented on resource-constrained portable/wearable devices with limited capacity of battery. To address this design challenge, a power-efficient CNN design through zero-gating processing elements (PEs) and partial-sum reuse centric dataflow is proposed in this paper. Unlike the existing works which either only consider the zeros in activation maps or use off-chip training process for on-chip computation reduction, a zero-gating PE design is proposed to avoid unnecessary on-chip computation by taking advantages of the large number of zeros in both the filter’s weights of pre-trained models and the activation maps. Furthermore, a partial-sum reuse centric dataflow is also proposed for off-chip DRAM access reduction. The evaluation results show that the overall power consumption of PE arrays with our proposal can be reduced by 37% and 14% at the cost of 8% and 1% area overhead when compared to the baseline PE design and the existing only-activation-gated design (i.e. that in Eyeriss), respectively. Moreover, the proposed method can achieve 35% and 47% DRAM access reduction with the corresponding 14% and 49% energy savings for AlexNet and VGG-16 when compared to that in Eyeriss.