An Energy-Efficient Deep Convolutional Neural Network Accelerator Featuring Conditional Computing and Low External Memory Access

An Energy-Efficient Deep Convolutional Neural Network Accelerator Featuring Conditional Computing and Low External Memory Access
复制标题

DOI:
10.1109/jssc.2020.3029235
复制
发表时间:
2020-10
影响因子:
5.4
通讯作者:
Minkyu Kim;Jae-sun Seo
Minkyu Kim;Jae-sun Seo
中科院分区:
工程技术1区
文献类型:
--
作者:
Minkyu Kim;Jae-sun Seo

文献摘要

被引文献

相似文献

凭借其在许多机器学习任务和应用中的算法成功,深度卷积神经网络(DCNN)已经在许多先前的工作中使用定制硬件实现。然而,这些工作并没有最大限度地利用条件/近似计算来消除CNN的冗余计算。本文介绍了一种DCNN加速器,其具有一种新颖的条件计算方案,该方案将精确级联(PC)与零跳过(ZS)协同结合起来。为了减少最大池操作之后的许多冗余卷积,我们提出了精度级联,其中输入特征被划分为许多低精度组,并且首先执行仅具有最高有效位(MSB)的近似卷积。基于这种近似计算,全精度卷积仅在找到的最大池化输出上执行。这样,按位卷积的总数可以减少$\sim 2\times $,ImageNet精度下降< 0.8%。PC提供了增加每个低精度组的稀疏性的额外好处,我们利用ZS来消除时钟周期和外部存储器访问。所提出的条件计算方案已经在40纳米原型芯片中用定制架构实现,其在0.6-V电源下实现了24.97 TOPS/W的峰值能效,并且使用VGG-16 CNN进行ImageNet分类时的外部存储器访问为0.0018访问/MAC,并且在0.9-V电源下实现了28.51 TOPS/W的峰值能效。为Flying Chair数据集提供FlowNet。
With its algorithmic success in many machine learning tasks and applications, deep convolutional neural networks (DCNNs) have been implemented with custom hardware in a number of prior works. However, such works have not exploited conditional/approximate computing to the utmost toward eliminating redundant computations of CNNs. This article presents a DCNN accelerator featuring a novel conditional computing scheme that synergistically combines precision cascading (PC) with zero skipping (ZS). To reduce many redundant convolutions that are followed by max-pooling operations, we propose precision cascading, where the input features are divided into a number of low-precision groups and approximate convolutions with only the most significant bits (MSBs) are performed first. Based on this approximate computation, the full-precision convolution is performed only on the maximum pooling output that is found. This way, the total number of bit-wise convolutions can be reduced by $\sim 2\times $ with < 0.8% degradation in ImageNet accuracy. PC provides the added benefit of increased sparsity per low-precision group, which we exploit with ZS to eliminate the clock cycles and external memory accesses. The proposed conditional computing scheme has been implemented with custom architecture in a 40-nm prototype chip, which achieves a peak energy efficiency of 24.97 TOPS/W at 0.6-V supply and a low external memory access of 0.0018 access/MAC with VGG-16 CNN for ImageNet classification and a peak energy efficiency of 28.51 TOPS/W at 0.9-V supply with FlowNet for Flying Chair data set.