Back Razor: Memory-Efficient Transfer Learning by Self-Sparsified Backpropagation

Back Razor: Memory-Efficient Transfer Learning by Self-Sparsified Backpropagation
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Ziyu Jiang;Xuxi Chen;Xueqin Huang;Xianzhi Du;Denny Zhou;Zhangyang Wang
Ziyu Jiang;Xuxi Chen;Xueqin Huang;Xianzhi Du;Denny Zhou;Zhangyang Wang
中科院分区:
其他
文献类型:
--
作者:
Ziyu Jiang;Xuxi Chen;Xueqin Huang;Xianzhi Du;Denny Zhou;Zhangyang Wang

文献摘要

相似文献

从大数据集上训练的模型到定制的下游任务的迁移学习已经被广泛应用,因为预训练的模型可以大大提高泛化能力。然而,预训练模型尺寸的增加也会导致下游传输的内存占用过大,使其无法用于个人设备。以前的工作认识到占用空间的瓶颈是激活,因此提出了各种解决方案,例如注入特定的生命模块。在这项工作中,我们提出了一种新的内存高效传输框架,称为Back Razor,它可以即插即用地应用于任何预训练的网络,而无需改变其架构。Back Razor的关键思想是不对称稀疏化:修剪存储用于反向传播的激活,同时保持正向激活的密集。它基于这样一种观察,即存储的激活(占内存占用的大部分)仅用于反向传播。这种不对称剪枝避免影响正向计算的精度,从而使更积极的剪枝成为可能。此外,我们对Back Razor的收敛速度进行了理论分析,表明在温和的条件下,我们的方法保持了与香草SGD相似的收敛速度。在卷积神经网络和视觉变形器的分类、密集预测和语言建模任务上进行的大量迁移学习实验表明,Back Razor可以产生高达97%的稀疏性,节省9.2倍的内存使用,而不会失去准确性。代码可从https://github.com/VITA-Group/BackRazor_Neurips22获得。
Transfer learning from the model trained on large datasets to customized down-stream tasks has been widely used as the pre-trained model can greatly boost the generalizability. However, the increasing sizes of pre-trained models also lead to a prohibitively large memory footprints for downstream transferring, making them unaffordable for personal devices. Previous work recognizes the bottleneck of the footprint to be the activation, and hence proposes various solutions such as injecting specific lite modules. In this work, we present a novel memory-efficient transfer framework called Back Razor , that can be plug-and-play applied to any pre-trained network without changing its architecture. The key idea of Back Razor is asymmetric sparsifying : pruning the activation stored for back-propagation, while keeping the forward activation dense. It is based on the observation that the stored activation, that dominates the memory footprint, is only needed for back-propagation. Such asymmetric pruning avoids affecting the precision of forward computation, thus making more aggressive pruning possible. Furthermore, we conduct the theoretical analysis for the convergence rate of Back Razor, showing that under mild conditions, our method retains the similar convergence rate as vanilla SGD. Extensive transfer learning experiments on both Convolutional Neural Networks and Vision Transformers with classification, dense prediction, and language modeling tasks show that Back Razor could yield up to 97% sparsity , saving 9.2x memory usage, without losing accuracy. The code is available at: https://github.com/VITA-Group/BackRazor_Neurips22 .