Towards Efficient Tensor Decomposition-Based DNN Model Compression with Optimization Framework

Towards Efficient Tensor Decomposition-Based DNN Model Compression with Optimization Framework
复制标题

DOI:
10.1109/cvpr46437.2021.01053
复制
发表时间:
2021-06
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Miao Yin;Yang Sui;Siyu Liao;Bo Yuan
Miao Yin;Yang Sui;Siyu Liao;Bo Yuan
中科院分区:
其他
文献类型:
--
作者:
Miao Yin;Yang Sui;Siyu Liao;Bo Yuan

文献摘要

被引文献

相似文献

高级张量分解,如张量训练(TT)和张量环(TR),已被广泛研究用于深度神经网络(DNN)模型压缩,特别是用于递归神经网络(RNN)。然而,使用TT/TR压缩卷积神经网络(CNN)总是遭受显着的准确性损失。在本文中,我们提出了一个系统的框架,张量分解为基础的模型压缩使用交替方向乘法(ADMM)。通过将基于TT分解的模型压缩公式化为一个具有张量秩约束的优化问题,我们利用ADMM技术以迭代的方式系统地解决这个优化问题。在此过程中,整个DNN模型以原始结构而不是TT格式进行训练,但逐渐享有所需的低张量秩特性。然后,我们将这个未压缩的模型分解为TT格式,并对其进行微调,最终获得高精度的TT格式DNN模型。我们的框架非常通用,适用于CNN和RNN,并且可以轻松修改以适应其他张量分解方法。我们在不同的DNN模型上评估了我们提出的框架,用于图像分类和视频识别任务。实验结果表明,我们的ADMM为基础的TT格式的模型表现出非常高的压缩性能和高精度。值得注意的是,在CIFAR-100上,2.3倍和2.4倍的压缩比,我们的模型分别比原始ResNet-20和ResNet-32高出1.96%和2.21%。对于在ImageNet上压缩ResNet-18,我们的模型实现了2.47× FLOPs的减少,而没有精度损失。
Advanced tensor decomposition, such as tensor train (TT) and tensor ring (TR), has been widely studied for deep neural network (DNN) model compression, especially for recurrent neural networks (RNNs). However, compressing convolutional neural networks (CNNs) using TT/TR always suffers significant accuracy loss. In this paper, we propose a systematic framework for tensor decomposition-based model compression using Alternating Direction Method of Multipliers (ADMM). By formulating TT decomposition-based model compression to an optimization problem with constraints on tensor ranks, we leverage ADMM technique to systemically solve this optimization problem in an iterative way. During this procedure, the entire DNN model is trained in the original structure instead of TT format, but gradually enjoys the desired low tensor rank characteristics. We then decompose this uncompressed model to TT format, and fine-tune it to finally obtain a high-accuracy TT-format DNN model. Our framework is very general, and it works for both CNNs and RNNs, and can be easily modified to fit other tensor decomposition approaches. We evaluate our proposed framework on different DNN models for image classification and video recognition tasks. Experimental results show that our ADMM-based TT-format models demonstrate very high compression performance with high accuracy. Notably, on CIFAR-100, with 2.3× and 2.4× compression ratios, our models have 1.96% and 2.21% higher top-1 accuracy than the original ResNet-20 and ResNet-32, respectively. For compressing ResNet-18 on ImageNet, our model achieves 2.47× FLOPs reduction without accuracy loss.