Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed

Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed
复制标题

对高维高斯混合物进行分类:核方法失败而神经网络成功的地方

DOI:
--
复制
发表时间:
2021
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
Lenka Zdeborov'a
Lenka Zdeborov'a
中科院分区:
--
文献类型:
--
作者:
Maria Refinetti;Sebastian Goldt;Florent Krzakala;Lenka Zdeborov'a

文献摘要

参考文献

被引文献

相似文献

最近的一系列理论工作表明,核方法可以很好地捕获具有一定初始化的神经网络的动态。同时进行的实证工作表明,核方法在某些图像分类任务上可以接近神经网络的性能。这些结果提出了一个问题:尽管神经网络更具表现力,但神经网络是否只有在内核也能成功学习的情况下才能成功学习。在这里,我们从理论上证明,仅具有少量隐藏神经元的两层神经网络(2LNN)可以在简单的高斯混合分类任务上击败核学习的性能。我们研究了样本数量与输入维度成线性比例的高维限制,结果表明,虽然小型 2LNN 在此任务上实现了近乎最佳的性能,但随机特征和核方法等惰性训练方法却无法做到这一点。我们的分析基于一组封闭方程的推导,这些方程跟踪 2LNN 的学习动态,从而可以提取网络的渐近性能作为信噪比和其他超参数的函数。我们最后说明了神经网络的过度参数化如何导致更快的收敛,但并没有提高其最终性能。
A recent series of theoretical works showed that the dynamics of neural networks with a certain initialisation are well-captured by kernel methods. Concurrent empirical work demonstrated that kernel methods can come close to the performance of neural networks on some image classification tasks. These results raise the question of whether neural networks only learn successfully if kernels also learn successfully, despite neural networks being more expressive. Here, we show theoretically that two-layer neural networks (2LNN) with only a few hidden neurons can beat the performance of kernel learning on a simple Gaussian mixture classification task. We study the high-dimensional limit where the number of samples is linearly proportional to the input dimension, and show that while small 2LNN achieve near-optimal performance on this task, lazy training approaches such as random features and kernel methods do not. Our analysis is based on the derivation of a closed set of equations that track the learning dynamics of the 2LNN and thus allow to extract the asymptotic performance of the network as a function of signal-to-noise ratio and other hyperparameters. We finally illustrate how over-parametrising the neural network leads to faster convergence, but does not improve its final performance.
DOI: --
发表时间: 2018-02
期刊: ArXiv
影响因子: --
作者:
A. G. Matthews;Mark Rowland;Jiri Hron;Richard E. Turner;Zoubin Ghahramani
通讯作者: A. G. Matthews;Mark Rowland;Jiri Hron;Richard E. Turner;Zoubin Ghahramani
DOI: 10.1088/1742-5468/ac3a81
发表时间: 2020-06
期刊: Journal of Statistical Mechanics: Theory and Experiment
影响因子: --
作者:
B. Ghorbani;Song Mei;Theodor Misiakiewicz;A. Montanari
通讯作者: B. Ghorbani;Song Mei;Theodor Misiakiewicz;A. Montanari
DOI: --
发表时间: 2020
期刊: Thirty-seventh International Conference on Machine Learning (ICML
影响因子: --
作者:
Mignacco, Francesca;Krzakala, Florent;Lu, Yue M;Zdeborová, Lenka
通讯作者: Zdeborová, Lenka
DOI: 10.1002/cpa.22008
发表时间: 2019-08
影响因子: 3
作者:
Song Mei;A. Montanari
通讯作者: Song Mei;A. Montanari