Spectral Bias Outside the Training Set for Deep Networks in the Kernel Regime

Spectral Bias Outside the Training Set for Deep Networks in the Kernel Regime
复制标题

DOI:
10.48550/arxiv.2206.02927
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Benjamin Bowman;Guido Montúfar
Benjamin Bowman;Guido Montúfar
中科院分区:
其他
文献类型:
--
作者:
Benjamin Bowman;Guido Montúfar

文献摘要

被引文献

相似文献

我们从无限宽度的理想核动力学和无限数据的理想核动力学出发,给出了在有限个样本上训练的有限宽度网络的轨迹在函数空间上的$L^2$差的量化界限。界限的一个含义是,网络偏向于不仅在训练集上而且在整个输入空间上学习神经正切核的顶部特征函数。这种偏向仅取决于模型体系结构和输入分布,因此不依赖于不需要位于内核的RKHS中的目标函数。该结果对于具有完全连通、卷积和残留层的深层结构是有效的。此外,为了获得直到停止时间的高概率界,宽度不需要与样本数以多项式增长。该证明利用了Fisher信息矩阵在初始化时的低有效秩属性,这意味着模型的有效维度很低(远远小于参数的数目)。我们的结论是,从Fisher信息矩阵的低有效秩处进行局部容量控制在理论上仍未得到充分的探索。
We provide quantitative bounds measuring the $L^2$ difference in function space between the trajectory of a finite-width network trained on finitely many samples from the idealized kernel dynamics of infinite width and infinite data. An implication of the bounds is that the network is biased to learn the top eigenfunctions of the Neural Tangent Kernel not just on the training set but over the entire input space. This bias depends on the model architecture and input distribution alone and thus does not depend on the target function which does not need to be in the RKHS of the kernel. The result is valid for deep architectures with fully connected, convolutional, and residual layers. Furthermore the width does not need to grow polynomially with the number of samples in order to obtain high probability bounds up to a stopping time. The proof exploits the low-effective-rank property of the Fisher Information Matrix at initialization, which implies a low effective dimension of the model (far smaller than the number of parameters). We conclude that local capacity control from the low effective rank of the Fisher Information Matrix is still underexplored theoretically.