Surrogate Lagrangian Relaxation: A Path to Retrain-Free Deep Neural Network Pruning

Surrogate Lagrangian Relaxation: A Path to Retrain-Free Deep Neural Network Pruning
复制标题

DOI:
10.1145/3624476
复制
发表时间:
2023-04
影响因子:
1.4
通讯作者:
Shangli Zhou;Mikhail A. Bragin;Lynn Pepin;Deniz Gurevin;Fei Miao;Caiwen Ding
Shangli Zhou;Mikhail A. Bragin;Lynn Pepin;Deniz Gurevin;Fei Miao;Caiwen Ding
中科院分区:
计算机科学4区
文献类型:
--
作者:
Shangli Zhou;Mikhail A. Bragin;Lynn Pepin;Deniz Gurevin;Fei Miao;Caiwen Ding

文献摘要

相似文献

网络修剪是一种广泛使用的技术,用于降低深度神经网络的计算成本和模型大小。然而,典型的三级流水线(即,训练、修剪和再训练(微调))显著增加了总训练时间。在这篇文章中,我们开发了一个系统的权重修剪优化方法的基础上代理拉格朗日松弛(SLR),这是量身定制的,以克服所造成的困难的离散性质的权重修剪问题。我们进一步证明,我们的方法确保快速收敛的模型压缩问题,并通过使用二次惩罚的SLR的收敛速度加快。与其他最先进的方法相比,在训练阶段通过SLR获得的模型参数更接近其最优值。我们使用CIFAR-10和ImageNet对图像分类任务进行了评估,其中包括最先进的基于多层感知器的网络,如MLP Mixer;基于注意力的网络,如Swin Transformer;以及基于卷积神经网络的模型,如VGG-16,ResNet-18,ResNet-50,ResNet-110和MobileNetV 2。我们还使用各种模型在COCO,KITTI基准和TuSimple车道检测数据集上评估对象检测和分割任务。实验结果表明,在相同的精度要求下,该方法比现有方法具有更高的压缩率,并且在相同的压缩率要求下,该方法也具有更高的精度.在分类任务下,我们的SLR方法在两个数据集上都能更快地收敛到所需的精度。在目标检测和分割任务下,SLR也能以2倍的速度收敛到所需的精度。此外,我们的SLR实现了高模型精度,即使在硬修剪阶段没有再训练,这减少了传统的三阶段修剪成两个阶段的过程。在有限的再训练周期预算下,我们的方法可以快速恢复模型的准确性。
Network pruning is a widely used technique to reduce computation cost and model size for deep neural networks. However, the typical three-stage pipeline (i.e., training, pruning, and retraining (fine-tuning)) significantly increases the overall training time. In this article, we develop a systematic weight-pruning optimization approach based on surrogate Lagrangian relaxation (SLR), which is tailored to overcome difficulties caused by the discrete nature of the weight-pruning problem. We further prove that our method ensures fast convergence of the model compression problem, and the convergence of the SLR is accelerated by using quadratic penalties. Model parameters obtained by SLR during the training phase are much closer to their optimal values as compared to those obtained by other state-of-the-art methods. We evaluate our method on image classification tasks using CIFAR-10 and ImageNet with state-of-the-art multi-layer perceptron based networks such as MLP-Mixer; attention-based networks such as Swin Transformer; and convolutional neural network based models such as VGG-16, ResNet-18, ResNet-50, ResNet-110, and MobileNetV2. We also evaluate object detection and segmentation tasks on COCO, the KITTI benchmark, and the TuSimple lane detection dataset using a variety of models. Experimental results demonstrate that our SLR-based weight-pruning optimization approach achieves a higher compression rate than state-of-the-art methods under the same accuracy requirement and also can achieve higher accuracy under the same compression rate requirement. Under classification tasks, our SLR approach converges to the desired accuracy × faster on both of the datasets. Under object detection and segmentation tasks, SLR also converges 2× faster to the desired accuracy. Further, our SLR achieves high model accuracy even at the hardpruning stage without retraining, which reduces the traditional three-stage pruning into a two-stage process. Given a limited budget of retraining epochs, our approach quickly recovers the model’s accuracy.