Modelling the influence of data structure on learning in neural networks

Modelling the influence of data structure on learning in neural networks
复制标题

建模数据结构对神经网络学习的影响

DOI:
--
复制
发表时间:
2019
期刊:
arXiv.org
影响因子:
--
通讯作者:
L. Zdeborová
L. Zdeborová
中科院分区:
--
文献类型:
--
作者:
Sebastian Goldt;M. Mézard;Florent Krzakala;L. Zdeborová

文献摘要

参考文献

被引文献

相似文献

理解使用基于随机梯度的方法训练的深度神经网络成功的原因,是新兴的深度学习理论的一个关键未解决问题。这些网络最为成功的数据类型,例如图像或语音序列,具有复杂的相关性。然而,大多数关于神经网络的理论工作并没有明确地对训练数据进行建模,或者假设每个数据样本的元素是从某个分解的概率分布中独立抽取的。因此,这些方法从构建上就对现实世界数据集的相关性结构及其对神经网络学习的影响视而不见。在此,我们引入一种用于结构化数据集的生成模型,我们称之为隐流形模型(HMM)。其思路是构建位于低维流形上的高维输入,其标签仅取决于它们在该流形内的位置,类似于生成对抗网络中的单层解码器或生成器。我们通过证明一个“高斯等效性质”(GEP),表明隐流形模型的学习可进行解析处理,并且我们使用GEP来展示如何通过一组积分 - 微分方程来捕捉使用一次通过随机梯度下降训练的两层神经网络的动态,这些方程可随时跟踪网络的性能。这使我们能够详细分析神经网络在训练过程中如何学习复杂度不断增加的函数,其性能如何取决于其规模,以及它如何受到诸如学习率或隐流形维度等参数的影响。
Understanding the reasons for the success of deep neural networks trained using stochastic gradient-based methods is a key open problem for the nascent theory of deep learning. The types of data where these networks are most successful, such as images or sequences of speech, are characterised by intricate correlations. Yet, most theoretical work on neural networks does not explicitly model training data, or assumes that elements of each data sample are drawn independently from some factorised probability distribution. These approaches are thus by construction blind to the correlation structure of real-world data sets and their impact on learning in neural networks. Here, we introduce a generative model for structured data sets that we call the hidden manifold model (HMM). The idea is to construct high-dimensional inputs that lie on a lower-dimensional manifold, with labels that depend only on their position within this manifold, akin to a single layer decoder or generator in a generative adversarial network. We demonstrate that learning of the hidden manifold model is amenable to an analytical treatment by proving a "Gaussian Equivalence Property" (GEP), and we use the GEP to show how the dynamics of two-layer neural networks trained using one-pass stochastic gradient descent is captured by a set of integro-differential equations that track the performance of the network at all times. This permits us to analyse in detail how a neural network learns functions of increasing complexity during training, how its performance depends on its size and how it is impacted by parameters such as the learning rate or the dimension of the hidden manifold.
DOI: 10.1002/cpa.22008
发表时间: 2019-08
影响因子: 3
作者:
Song Mei;A. Montanari
通讯作者: Song Mei;A. Montanari