AdaPrune: An Accelerator-Aware Pruning Technique for Sustainable CNN Accelerators

AdaPrune: An Accelerator-Aware Pruning Technique for Sustainable CNN Accelerators
复制标题

DOI:
10.1109/tsusc.2021.3060690
复制
发表时间:
2022-01-01
影响因子:
3.9
通讯作者:
Louri, Ahmed
Louri, Ahmed
中科院分区:
计算机科学2区
文献类型:
--
作者:
Li, Jiajun;Louri, Ahmed

文献摘要

被引文献

相似文献

卷积神经网络(CNN)加速器已经从云到边缘场景取得了巨大成功。然而,考虑到更大更深的神经网络模型的趋势,有效地处理这些CNN仍然是一个具有挑战性的问题,特别是在能量预算有限的边缘设备上。因此,降低能耗对于可持续的CNN加速器至关重要。在本文中,我们提出了AdaPrune,这是一种新的修剪技术,可以减少模型大小和计算量,以实现CNN加速器的性能提高和节能。与以前牺牲计算规律性或准确性的修剪技术不同,AdaPrune通过为底层加速器定制CNN修剪来保持两者,以最大限度地利用稀疏性优势。AdaPrune由两种技术组成:输入通道组修剪和输出通道组修剪。通过分析稀疏CNN加速器的权重提取模式,AdaPrune自适应地在两种技术之间切换,以保证零在每个提取组中均匀分布。在这样做的过程中,修剪后的网络结构为底层加速器保留了定制的计算规律,从而提高了性能和能效。我们在三个稀疏CNN加速器上使用不同的空间平铺策略评估AdaPrune。实验结果表明,AdaPrune实现了高达1.6倍的性能加速,和1.5倍的节能相比,非结构化修剪。
Convolutional neural network (CNN) accelerators have achieved great success from cloud to edge scenarios. However, given the trend towards even larger and deeper neural network models, it remains a challenging problem to efficiently process these CNNs especially on edge devices with limited energy budget. Accordingly, reducing the energy consumption is of paramount importance for sustainable CNN accelerators. In this paper, we propose AdaPrune, a novel pruning technique that reduces model size and computation to achieve performance improvement and energy savings for CNN accelerators. Unlike previous pruning techniques that sacrifice either computational regularity or accuracy, AdaPrune maintains both by customizing CNN pruning for the underlying accelerators to maximally leverage the sparsity benefits. AdaPrune consists of two techniques: input channel group pruning and output channel group pruning. By analyzing the weight fetching patterns of sparse CNN accelerators, AdaPrune adaptively switches between the two techniques to guarantee that the zeros are evenly distributed in each fetching group. In doing so, the pruned network structure preserves customized computational regularity for the underlying accelerators, thereby boosting the performance and energy efficiency. We evaluate AdaPrune on three sparse CNN accelerators with different spatial tiling strategies. The experimental results show that AdaPrune achieves up to 1.6x performance speedup, and 1.5x energy savings compared to unstructured pruning.