Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural Networks

Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural Networks
复制标题

DOI:
10.1609/aaai.v35i8.16859
复制
发表时间:
2020-12
期刊:
--
影响因子:
--
通讯作者:
Xiangyu Chang;Yingcong Li;Samet Oymak;Christos Thrampoulidis
Xiangyu Chang;Yingcong Li;Samet Oymak;Christos Thrampoulidis
中科院分区:
其他
文献类型:
--
作者:
Xiangyu Chang;Yingcong Li;Samet Oymak;Christos Thrampoulidis

文献摘要

相似文献

深度网络通常使用比训练数据集的大小更多的参数进行训练。最近的经验证据表明,过度参数化的实践不仅有利于训练大型模型,而且还有助于(可能与直觉相反)构建轻量级模型。具体来说,它表明过度参数化有利于模型修剪/稀疏化。本文从理论上刻画了过参数化状态下模型剪枝的高维渐近性,从而阐明了这些经验发现。提出的理论解决了以下核心问题:“应该从一开始训练一个小模型,还是先训练一个大模型,然后进行修剪?”我们通过分析来确定这样的制度,即使最具信息量的特征的位置是已知的,我们最好是拟合一个大模型,然后进行修剪,而不是简单地用已知的信息量特征进行训练。这导致了稀疏模型训练中的一种新的双重下降:在保持目标稀疏性的同时,增加原始模型,随着超过过参数化阈值而提高测试精度。我们的分析进一步揭示了将再训练与特征相关性联系起来的好处。我们发现上述现象已经存在于线性和随机特征模型中。我们的技术方法提高了高维分析的工具集,并精确地表征了过参数化最小二乘的渐近分布。通过分析研究简单模型所获得的直觉在神经网络上得到了数值验证。
Deep networks are typically trained with many more parameters than the size of the training dataset. Recent empirical evidence indicates that the practice of overparameterization not only benefits training large models, but also assists – perhaps counterintuitively – building lightweight models. Specifically, it suggests that overparameterization benefits model pruning / sparsification. This paper sheds light on these empirical findings by theoretically characterizing the high-dimensional asymptotics of model pruning in the overparameterized regime. The theory presented addresses the following core question: ``should one train a small model from the beginning, or first train a large model and then prune?''. We analytically identify regimes in which, even if the location of the most informative features is known, we are better off fitting a large model and then pruning rather than simply training with the known informative features. This leads to a new double descent in the training of sparse models: growing the original model, while preserving the target sparsity, improves the test accuracy as one moves beyond the overparameterization threshold. Our analysis further reveals the benefit of retraining by relating it to feature correlations. We find that the above phenomena are already present in linear and random-features models. Our technical approach advances the toolset of high-dimensional analysis and precisely characterizes the asymptotic distribution of over-parameterized least-squares. The intuition gained by analytically studying simpler models is numerically verified on neural networks.