EDP: An Efficient Decomposition and Pruning Scheme for Convolutional Neural Network Compression

EDP: An Efficient Decomposition and Pruning Scheme for Convolutional Neural Network Compression
复制标题

EDP​​:一种高效的卷积神经网络压缩分解和剪枝方案

DOI:
10.1109/tnnls.2020.3018177
复制
发表时间:
2020-11
影响因子:
10.4
通讯作者:
Stephen Maybank
Stephen Maybank
中科院分区:
计算机科学1区
文献类型:
--
作者:
Xiaofeng Ruan;Yufan Liu;Chunfeng Yuan;Bing Li;Weiming Hu;Yangxi Li;Stephen Maybank

文献摘要

参考文献

相似文献

模型压缩方法近年来变得流行,其目的是减轻现实世界应用中深度神经网络(DNN)的沉重负载。然而,大多数现有的压缩方法有两个局限性:1)它们通常采用繁琐的过程,包括预训练,稀疏约束训练,修剪/分解和微调。此外,最后三个阶段通常迭代多次。2)模型是在显式稀疏或低秩假设下预训练的,这很难保证广泛的适用性。在这篇文章中,我们提出了一个有效的分解和修剪(EDP)计划,通过构建一个压缩感知块,可以自动最小化的权重矩阵的秩和识别冗余通道。具体来说,我们通过将一个网络层分解为两层来嵌入压缩感知块:新的权重矩阵层和系数矩阵层。通过在系数矩阵上施加正则化器,新的权重矩阵学习成为低秩基权重,并且其对应的信道变得稀疏。以这种方式,所提出的压缩感知块同时实现低秩分解和通道修剪,只有一个单一的数据驱动的训练阶段。此外,通过一种新的剪枝合并模块,对结构网络进行了进一步的压缩和优化,该模块对冗余通道进行剪枝,并对冗余分解层进行合并。在不同的数据集和网络上的实验结果(17个竞争对手)表明,所提出的EDP实现了高压缩比与可接受的准确性下降,并优于国家的最先进的压缩率,准确性,推理时间和运行时内存。
Model compression methods have become popular in recent years, which aim to alleviate the heavy load of deep neural networks (DNNs) in real-world applications. However, most of the existing compression methods have two limitations: 1) they usually adopt a cumbersome process, including pretraining, training with a sparsity constraint, pruning/decomposition, and fine-tuning. Moreover, the last three stages are usually iterated multiple times. 2) The models are pretrained under explicit sparsity or low-rank assumptions, which are difficult to guarantee wide appropriateness. In this article, we propose an efficient decomposition and pruning (EDP) scheme via constructing a compressed-aware block that can automatically minimize the rank of the weight matrix and identify the redundant channels. Specifically, we embed the compressed-aware block by decomposing one network layer into two layers: a new weight matrix layer and a coefficient matrix layer. By imposing regularizers on the coefficient matrix, the new weight matrix learns to become a low-rank basis weight, and its corresponding channels become sparse. In this way, the proposed compressed-aware block simultaneously achieves low-rank decomposition and channel pruning by only one single data-driven training stage. Moreover, the network of architecture is further compressed and optimized by a novel Pruning & Merging (PM) module which prunes redundant channels and merges redundant decomposed layers. Experimental results (17 competitors) on different data sets and networks demonstrate that the proposed EDP achieves a high compression ratio with acceptable accuracy degradation and outperforms state-of-the-arts on compression rate, accuracy, inference time, and run-time memory.
DOI: --
发表时间: 2018-02
期刊: --
影响因子: --
作者:
Dejiao Zhang;Haozhu Wang;Mário A. T. Figueiredo;L. Balzano
通讯作者: Dejiao Zhang;Haozhu Wang;Mário A. T. Figueiredo;L. Balzano
DOI: --
发表时间: 2016-11
期刊: ArXiv
影响因子: --
作者:
Chenzhuo Zhu;Song Han;Huizi Mao;W. Dally
通讯作者: Chenzhuo Zhu;Song Han;Huizi Mao;W. Dally
DOI: --
发表时间: 2018-09
期刊: ArXiv
影响因子: --
作者:
Namhoon Lee;Thalaiyasingam Ajanthan;Philip H. S. Torr
通讯作者: Namhoon Lee;Thalaiyasingam Ajanthan;Philip H. S. Torr
DOI: --
发表时间: 2009
期刊: --
影响因子: --
作者:
A. Krizhevsky
通讯作者: A. Krizhevsky
DOI: --
发表时间: 2019-05
期刊: ArXiv
影响因子: --
作者:
Xiaohan Ding;Guiguang Ding;Yuchen Guo;J. Han;C. Yan
通讯作者: Xiaohan Ding;Guiguang Ding;Yuchen Guo;J. Han;C. Yan