EDP: An Efficient Decomposition and Pruning Scheme for Convolutional Neural Network Compression
EDP: An Efficient Decomposition and Pruning Scheme for Convolutional Neural Network Compression
复制标题
EDP:一种高效的卷积神经网络压缩分解和剪枝方案
DOI:
10.1109/tnnls.2020.3018177
复制
发表时间:
2020-11
影响因子:
10.4
通讯作者:
Stephen Maybank
中科院分区:
文献类型:
--
作者:
Xiaofeng Ruan;Yufan Liu;Chunfeng Yuan;Bing Li;Weiming Hu;Yangxi Li;Stephen Maybank
Model compression methods have become popular in recent years, which aim to alleviate the heavy load of deep neural networks (DNNs) in real-world applications. However, most of the existing compression methods have two limitations: 1) they usually adopt a cumbersome process, including pretraining, training with a sparsity constraint, pruning/decomposition, and fine-tuning. Moreover, the last three stages are usually iterated multiple times. 2) The models are pretrained under explicit sparsity or low-rank assumptions, which are difficult to guarantee wide appropriateness. In this article, we propose an efficient decomposition and pruning (EDP) scheme via constructing a compressed-aware block that can automatically minimize the rank of the weight matrix and identify the redundant channels. Specifically, we embed the compressed-aware block by decomposing one network layer into two layers: a new weight matrix layer and a coefficient matrix layer. By imposing regularizers on the coefficient matrix, the new weight matrix learns to become a low-rank basis weight, and its corresponding channels become sparse. In this way, the proposed compressed-aware block simultaneously achieves low-rank decomposition and channel pruning by only one single data-driven training stage. Moreover, the network of architecture is further compressed and optimized by a novel Pruning & Merging (PM) module which prunes redundant channels and merges redundant decomposed layers. Experimental results (17 competitors) on different data sets and networks demonstrate that the proposed EDP achieves a high compression ratio with acceptable accuracy degradation and outperforms state-of-the-arts on compression rate, accuracy, inference time, and run-time memory.
登录
查看更多内容
DOI:
--
发表时间:
2018-02
期刊:
--
影响因子:
--
作者:
Dejiao Zhang;Haozhu Wang;Mário A. T. Figueiredo;L. Balzano
通讯作者:
Dejiao Zhang;Haozhu Wang;Mário A. T. Figueiredo;L. Balzano
DOI:
--
发表时间:
2016-11
期刊:
ArXiv
影响因子:
--
作者:
Chenzhuo Zhu;Song Han;Huizi Mao;W. Dally
通讯作者:
Chenzhuo Zhu;Song Han;Huizi Mao;W. Dally
DOI:
--
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
作者:
Namhoon Lee;Thalaiyasingam Ajanthan;Philip H. S. Torr
通讯作者:
Namhoon Lee;Thalaiyasingam Ajanthan;Philip H. S. Torr
DOI:
--
发表时间:
2009
期刊:
--
影响因子:
--
作者:
A. Krizhevsky
通讯作者:
A. Krizhevsky
DOI:
--
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
作者:
Xiaohan Ding;Guiguang Ding;Yuchen Guo;J. Han;C. Yan
通讯作者:
Xiaohan Ding;Guiguang Ding;Yuchen Guo;J. Han;C. Yan