On the Neural Tangent Kernel Analysis of Randomly Pruned Neural Networks

On the Neural Tangent Kernel Analysis of Randomly Pruned Neural Networks
复制标题

DOI:
--
复制
发表时间:
2022-03
期刊:
--
影响因子:
--
通讯作者:
Hongru Yang;Zhangyang Wang
Hongru Yang;Zhangyang Wang
中科院分区:
其他
文献类型:
--
作者:
Hongru Yang;Zhangyang Wang

文献摘要

相似文献

在理论和实践的双重激励下,我们研究了随机剪枝对神经网络的神经切核的影响。特别是,这项工作建立了完全连通神经网络与其随机剪枝版本之间的NTK的等价性。在两种情况下建立了等价性。第一个主要结果研究了无穷宽的渐近性。结果表明,在给定剪枝概率的情况下,对于权值在初始化时随机剪枝的全连通神经网络,随着每一层的宽度依次增长到无穷大,被剪枝神经网络的NTK以一定的比例收敛到原始网络的极限NTK。如果在修剪后适当地重新缩放网络权重,则可以删除这种额外的缩放。第二个主要结果考虑了有限宽度的情形。结果表明,为了保证NTK接近极限,当NTK到极限的间隙减小到零时,宽度对稀疏性参数的依赖关系是渐近线性的。此外,如果剪枝概率设置为零(即不剪枝),则所需宽度的界与以前工作中的完全连通神经网络的界匹配,直到对数因子。为了证明这一结果,需要发展一种新的网络结构分析方法,我们称之为TECHITT(掩码诱导伪网络)。提供了实验来评估我们的结果。
Motivated by both theory and practice, we study how random pruning of the weights affects a neural network's neural tangent kernel (NTK). In particular, this work establishes an equivalence of the NTKs between a fully-connected neural network and its randomly pruned version. The equivalence is established under two cases. The first main result studies the infinite-width asymptotic. It is shown that given a pruning probability, for fully-connected neural networks with the weights randomly pruned at the initialization, as the width of each layer grows to infinity sequentially, the NTK of the pruned neural network converges to the limiting NTK of the original network with some extra scaling. If the network weights are rescaled appropriately after pruning, this extra scaling can be removed. The second main result considers the finite-width case. It is shown that to ensure the NTK's closeness to the limit, the dependence of width on the sparsity parameter is asymptotically linear, as the NTK's gap to its limit goes down to zero. Moreover, if the pruning probability is set to zero (i.e., no pruning), the bound on the required width matches the bound for fully-connected neural networks in previous works up to logarithmic factors. The proof of this result requires developing a novel analysis of a network structure which we called \textit{mask-induced pseudo-networks}. Experiments are provided to evaluate our results.