Non-Structured DNN Weight Pruning--Is It Beneficial in Any Platform?

Non-Structured DNN Weight Pruning--Is It Beneficial in Any Platform?
复制标题

DOI:
10.1109/tnnls.2021.3063265
复制
发表时间:
2021-03-18
影响因子:
10.4
通讯作者:
Wang, Yanzhi
Wang, Yanzhi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ma, Xiaolong;Lin, Sheng;Wang, Yanzhi

文献摘要

被引文献

相似文献

大型深度神经网络(DNN)模型是能源效率的关键挑战,因为片外DRAM访问的能耗比算术或SRAM操作高得多。它通过两种主要途径推动了模型压缩的深入研究。权重剪枝利用了权重数目中的冗余,并且可以以非结构化的方式执行,非结构化的方式具有较高的灵活性和剪枝率,但由于不规则的权重而导致索引访问,或者结构化的方式,其以较低的剪枝率保留完整的矩阵结构。权重量化利用了权重中的比特数中的冗余。与修剪相比,量化对硬件更加友好,并且已成为FPGA和ASIC实现的“必做”步骤。因此,对修剪效果的任何评估都应该放在量化的基础上。关键的悬而未决的问题是,对于量化,什么样的修剪(非结构化与结构化)是最有益的?这个问题是根本的,因为答案将决定我们真正应该关注的设计方面,以避免某些优化的回报递减。这篇文章首次为这个问题提供了一个明确的答案。首先,我们通过扩展和增强最近提出的联合权重剪枝和量化框架ADMM-NN来构建ADMM-NN-S,该框架具有结构化剪枝、动态ADMM调整以及掩蔽映射和再训练的算法支持。其次,我们开发了一种方法,从存储效率和计算效率两个方面对非结构化剪枝和结构化剪枝进行公平和基本的比较。我们的结果表明,ADMM-NN-S的性能一直优于现有技术:1)它在LeNet-5、AlexNet和ResNet-50上分别实现了348倍、36倍和8倍的总权重剪枝,精度损失(几乎)为零;2)我们证明了第一个完全二值化(对所有层)的DNN在许多情况下都可以在精度上无损。这些结果为我们的研究提供了强有力的基线和可信度。基于所提出的比较框架,在相同的精度和量化条件下,结果表明非结构化剪枝在存储和计算效率方面都不具有竞争力。因此,我们得出结论,与非结构化修剪相比,结构化修剪具有更大的潜力。我们鼓励社会各界集中研究结构稀疏的DNN推理加速。
Large deep neural network (DNN) models pose the key challenge to energy efficiency due to the significantly higher energy consumption of off-chip DRAM accesses than arithmetic or SRAM operations. It motivates the intensive research on model compression with two main approaches. Weight pruning leverages the redundancy in the number of weights and can be performed in a non-structured, which has higher flexibility and pruning rate but incurs index accesses due to irregular weights, or structured manner, which preserves the full matrix structure with a lower pruning rate. Weight quantization leverages the redundancy in the number of bits in weights. Compared to pruning, quantization is much more hardware-friendly and has become a ``must-do'' step for FPGA and ASIC implementations. Thus, any evaluation of the effectiveness of pruning should be on top of quantization. The key open question is, with quantization, what kind of pruning (non-structured versus structured) is most beneficial? This question is fundamental because the answer will determine the design aspects that we should really focus on to avoid the diminishing return of certain optimizations. This article provides a definitive answer to the question for the first time. First, we build ADMM-NN-S by extending and enhancing ADMM-NN, a recently proposed joint weight pruning and quantization framework, with the algorithmic supports for structured pruning, dynamic ADMM regulation, and masked mapping and retraining. Second, we develop a methodology for fair and fundamental comparison of non-structured and structured pruning in terms of both storage and computation efficiency. Our results show that ADMM-NN-S consistently outperforms the prior art: 1) it achieves 348x, 36x, and 8x overall weight pruning on LeNet-5, AlexNet, and ResNet-50, respectively, with (almost) zero accuracy loss and 2) we demonstrate the first fully binarized (for all layers) DNNs can be lossless in accuracy in many cases. These results provide a strong baseline and credibility of our study. Based on the proposed comparison framework, with the same accuracy and quantization, the results show that non-structured pruning is not competitive in terms of both storage and computation efficiency. Thus, we conclude that structured pruning has a greater potential compared to non-structured pruning. We encourage the community to focus on studying the DNN inference acceleration with structured sparsity.