AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates

AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates
复制标题

DOI:
10.1609/aaai.v34i04.5924
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Ning Liu;Xiaolong Ma;Zhiyuan Xu;Yanzhi Wang;Jian Tang;Jieping Ye
Ning Liu;Xiaolong Ma;Zhiyuan Xu;Yanzhi Wang;Jian Tang;Jieping Ye
中科院分区:
其他
文献类型:
--
作者:
Ning Liu;Xiaolong Ma;Zhiyuan Xu;Yanzhi Wang;Jian Tang;Jieping Ye

文献摘要

被引文献

相似文献

结构化权重剪枝是DNN的代表性模型压缩技术,用于减少存储和计算需求并加速推理。由于大量的灵活的超参数,一个自动的超参数确定过程是必要的。本文提出了一种自动结构化剪枝框架AutoCompress,该框架具有以下关键性能改进:(i)有效地将结构化剪枝方案的组合纳入自动过程;(ii)采用最先进的基于ADMM的结构化权重剪枝作为核心算法,并提出了一种创新的额外纯化步骤,以进一步降低权重而不损失准确性;以及(iii)开发有效的启发式搜索方法,并通过基于经验的引导搜索进行增强,以取代先前与目标修剪问题存在潜在不兼容性的深度强化学习技术。在CIFAR-10和ImageNet数据集上进行的大量实验表明,AutoCompress是在权重和FLOP数量上实现超高剪枝率的关键。例如,在相同的精度下,AutoCompress在修剪率方面比之前的自动模型压缩工作高出33倍(实际参数计数减少120倍)。在智能手机上的实际测量中,从AutoCompress框架中观察到了显著的推理加速。我们在匿名链接http://bit.ly/2VZ63dS上发布了这项工作的模型。
Structured weight pruning is a representative model compression technique of DNNs to reduce the storage and computation requirements and accelerate inference. An automatic hyperparameter determination process is necessary due to the large number of flexible hyperparameters. This work proposes AutoCompress, an automatic structured pruning framework with the following key performance improvements: (i) effectively incorporate the combination of structured pruning schemes in the automatic process; (ii) adopt the state-of-art ADMM-based structured weight pruning as the core algorithm, and propose an innovative additional purification step for further weight reduction without accuracy loss; and (iii) develop effective heuristic search method enhanced by experience-based guided search, replacing the prior deep reinforcement learning technique which has underlying incompatibility with the target pruning problem. Extensive experiments on CIFAR-10 and ImageNet datasets demonstrate that AutoCompress is the key to achieve ultra-high pruning rates on the number of weights and FLOPs that cannot be achieved before. As an example, AutoCompress outperforms the prior work on automatic model compression by up to 33× in pruning rate (120× reduction in the actual parameter count) under the same accuracy. Significant inference speedup has been observed from the AutoCompress framework on actual measurements on smartphone. We release models of this work at anonymous link: http://bit.ly/2VZ63dS.