When do neural networks outperform kernel methods?

When do neural networks outperform kernel methods?
复制标题

DOI:
10.1088/1742-5468/ac3a81
复制
发表时间:
2020-06
期刊:
Journal of Statistical Mechanics: Theory and Experiment
影响因子:
--
通讯作者:
B. Ghorbani;Song Mei;Theodor Misiakiewicz;A. Montanari
B. Ghorbani;Song Mei;Theodor Misiakiewicz;A. Montanari
中科院分区:
其他
文献类型:
--
作者:
B. Ghorbani;Song Mei;Theodor Misiakiewicz;A. Montanari

文献摘要

被引文献

相似文献

对于随机梯度下降(SGD)的初始化的一定尺度,宽神经网络(NN)已被证明是很好的近似再生核希尔伯特空间(RKHS)方法。最近的实证研究表明,对于某些分类任务,RKHS方法可以取代NN,而不会造成很大的性能损失。另一方面,众所周知,两层NN比RKHS编码更丰富的平滑度类,并且我们知道一些特殊示例,其中经过SGD训练的NN可证明优于RKHS。即使在宽网络限制中,对于初始化的不同缩放,也是如此。我们如何调和上述主张?NN在哪些任务上优于RKHS?如果协变量几乎是各向同性的,则RKHS方法会遭受维数灾难,而NN可以通过学习最佳低维表示来克服它。在这里,我们表明,这种灾难的维数变得温和,如果协变量显示相同的低维结构的目标函数,我们精确地描述了这种权衡。在这些结果的基础上,我们提出了尖峰协变量模型,可以在一个统一的框架中捕获在早期工作中观察到的行为。我们假设这样一个潜在的低维结构是存在于图像分类。我们在数值上测试了这一假设,表明训练分布的特定扰动降低RKHS方法的性能比NN更显着。
For a certain scaling of the initialization of stochastic gradient descent (SGD), wide neural networks (NN) have been shown to be well approximated by reproducing kernel Hilbert space (RKHS) methods. Recent empirical work showed that, for some classification tasks, RKHS methods can replace NNs without a large loss in performance. On the other hand, two-layers NNs are known to encode richer smoothness classes than RKHS and we know of special examples for which SGD-trained NN provably outperform RKHS. This is true even in the wide network limit, for a different scaling of the initialization. How can we reconcile the above claims? For which tasks do NNs outperform RKHS? If covariates are nearly isotropic, RKHS methods suffer from the curse of dimensionality, while NNs can overcome it by learning the best low-dimensional representation. Here we show that this curse of dimensionality becomes milder if the covariates display the same low-dimensional structure as the target function, and we precisely characterize this tradeoff. Building on these results, we present the spiked covariates model that can capture in a unified framework both behaviors observed in earlier work. We hypothesize that such a latent low-dimensional structure is present in image classification. We test numerically this hypothesis by showing that specific perturbations of the training distribution degrade the performances of RKHS methods much more significantly than NNs.