Advancing Model Pruning via Bi-level Optimization

Advancing Model Pruning via Bi-level Optimization
复制标题

DOI:
10.48550/arxiv.2210.04092
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yihua Zhang;Yuguang Yao;Parikshit Ram;Pu Zhao;Tianlong Chen;Min-Fong Hong;Yanzhi Wang;Sijia Liu-Siji
Yihua Zhang;Yuguang Yao;Parikshit Ram;Pu Zhao;Tianlong Chen;Min-Fong Hong;Yanzhi Wang;Sijia Liu-Siji
中科院分区:
其他
文献类型:
--
作者:
Yihua Zhang;Yuguang Yao;Parikshit Ram;Pu Zhao;Tianlong Chen;Min-Fong Hong;Yanzhi Wang;Sijia Liu-Siji

文献摘要

相似文献

实际应用中的部署约束需要对大规模深度学习模型进行修剪,即提高其权值稀疏性。正如彩票假设(LTH)所示,剪枝也有可能提高它们的泛化能力。在LTH的核心,迭代幅度修剪(IMP)是成功找到“中奖票”的主要修剪方法。然而,随着目标剪枝比的增加,IMP的计算成本会过高。为了减少计算开销,人们开发了各种高效的“一次性”剪枝方法,但这些方法通常无法像IMP那样找到中奖彩票。这就提出了如何缩小剪枝精度和剪枝效率之间的差距的问题。为了解决这个问题,我们追求模型修剪的算法进步。具体地说,我们从一个新颖的观点——双级优化(BLO)来阐述修剪问题。我们证明了BLO解释为有效实现IMP中使用的修剪-再训练学习范式提供了技术基础优化基础。我们还表明,所提出的面向双级优化的修剪方法(称为BiP)是一类具有双线性问题结构的特殊BLO问题。利用这种双线性,我们从理论上证明了BiP可以像一阶优化一样容易地求解,从而继承了计算效率。通过对5种模型架构和4个数据集进行结构化和非结构化剪枝的大量实验,我们证明了在大多数情况下,BiP比IMP能找到更好的中奖票,并且在计算上与一次性剪枝方案一样高效,在相同的模型精度和稀疏度水平下,比IMP的速度提高了2-7倍。
The deployment constraints in practical applications necessitate the pruning of large-scale deep learning models, i.e., promoting their weight sparsity. As illustrated by the Lottery Ticket Hypothesis (LTH), pruning also has the potential of improving their generalization ability. At the core of LTH, iterative magnitude pruning (IMP) is the predominant pruning method to successfully find 'winning tickets'. Yet, the computation cost of IMP grows prohibitively as the targeted pruning ratio increases. To reduce the computation overhead, various efficient 'one-shot' pruning methods have been developed, but these schemes are usually unable to find winning tickets as good as IMP. This raises the question of how to close the gap between pruning accuracy and pruning efficiency? To tackle it, we pursue the algorithmic advancement of model pruning. Specifically, we formulate the pruning problem from a fresh and novel viewpoint, bi-level optimization (BLO). We show that the BLO interpretation provides a technically-grounded optimization base for an efficient implementation of the pruning-retraining learning paradigm used in IMP. We also show that the proposed bi-level optimization-oriented pruning method (termed BiP) is a special class of BLO problems with a bi-linear problem structure. By leveraging such bi-linearity, we theoretically show that BiP can be solved as easily as first-order optimization, thus inheriting the computation efficiency. Through extensive experiments on both structured and unstructured pruning with 5 model architectures and 4 data sets, we demonstrate that BiP can find better winning tickets than IMP in most cases, and is computationally as efficient as the one-shot pruning schemes, demonstrating 2-7 times speedup over IMP for the same level of model accuracy and sparsity.