Finite Depth and Width Corrections to the Neural Tangent Kernel

Finite Depth and Width Corrections to the Neural Tangent Kernel
复制标题

DOI:
--
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
B. Hanin;M. Nica
B. Hanin;M. Nica
中科院分区:
其他
文献类型:
--
作者:
B. Hanin;M. Nica

文献摘要

被引文献

相似文献

我们证明在有限的深度和宽度下,在随机初始化的relu网络中的神经切线内核(NTK)的平均值和方差。标准偏差在网络深度与宽度的比率上是指数级。因此,即使在无限的过度参数化的极限中,如果同时倾向于无穷大,则NTK也不是确定性的。此外,我们证明,对于如此深厚的网络,NTK在训练过程中具有非平凡的演变,这表明其首次SGD更新的平均值在网络深度与宽度的比率上也是指数的。这与深度固定且网络宽度非常大的状态形成鲜明对比。我们的结果表明,与相对较浅和宽阔的网络不同,即使在所谓的懒惰训练制度中,深层和宽的relu网络也能够学习数据依赖性功能。
We prove the precise scaling, at finite depth and width, for the mean and variance of the neural tangent kernel (NTK) in a randomly initialized ReLU network. The standard deviation is exponential in the ratio of network depth to width. Thus, even in the limit of infinite overparameterization, the NTK is not deterministic if depth and width simultaneously tend to infinity. Moreover, we prove that for such deep and wide networks, the NTK has a non-trivial evolution during training by showing that the mean of its first SGD update is also exponential in the ratio of network depth to width. This is sharp contrast to the regime where depth is fixed and network width is very large. Our results suggest that, unlike relatively shallow and wide networks, deep and wide ReLU networks are capable of learning data-dependent features even in the so-called lazy training regime.