Deep Networks and the Multiple Manifold Problem

Deep Networks and the Multiple Manifold Problem
复制标题

DOI:
--
复制
发表时间:
2020-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Sam Buchanan;D. Gilboa;John Wright
Sam Buchanan;D. Gilboa;John Wright
中科院分区:
其他
文献类型:
--
作者:
Sam Buchanan;D. Gilboa;John Wright

文献摘要

相似文献

我们研究多流形问题,这是一个基于机器视觉应用建模的二分类任务,其中一个深度全连接神经网络被训练用于分离单位球面上的两个低维子流形。我们对一维情况进行了分析,针对一种简单的流形配置证明,当网络深度\(L\)相对于数据的某些几何和统计特性较大时,网络宽度\(n\)作为\(L\)的足够大的多项式增长,并且来自流形的独立同分布样本数量是\(L\)的多项式,随机初始化的梯度下降以高概率快速学会完美地对两个流形进行分类。我们的分析在一个具有实际动机的模型问题背景下展示了深度和宽度的具体益处:深度作为一种拟合资源,更大的深度对应更平滑的网络,能够更容易地分离类别流形,而宽度作为一种统计资源,使随机初始化的网络及其梯度能够集中。论证围绕神经正切核及其在过参数化神经网络训练的非渐近分析中的作用展开;在这方面的文献中,我们为深度全连接网络的神经正切核贡献了基本最优的集中速率,要求宽度\(n\gtrsim L\,\mathrm{poly}(d_0)\),以在单位球面\(\mathbb{S}^{n_0 - 1}\)的\(d_0\)维子流形上实现初始核的均匀集中,并且提供了一个非渐近框架,用于确定在神经正切核(NTK)机制下使用结构化数据训练的网络的泛化能力。证明大量使用鞅集中来最优地处理初始随机网络各层之间的统计相关性。这种方法应该有助于为其他网络架构建立类似的结果。
We study the multiple manifold problem, a binary classification task modeled on applications in machine vision, in which a deep fully-connected neural network is trained to separate two low-dimensional submanifolds of the unit sphere. We provide an analysis of the one-dimensional case, proving for a simple manifold configuration that when the network depth $L$ is large relative to certain geometric and statistical properties of the data, the network width $n$ grows as a sufficiently large polynomial in $L$, and the number of i.i.d. samples from the manifolds is polynomial in $L$, randomly-initialized gradient descent rapidly learns to classify the two manifolds perfectly with high probability. Our analysis demonstrates concrete benefits of depth and width in the context of a practically-motivated model problem: the depth acts as a fitting resource, with larger depths corresponding to smoother networks that can more readily separate the class manifolds, and the width acts as a statistical resource, enabling concentration of the randomly-initialized network and its gradients. The argument centers around the neural tangent kernel and its role in the nonasymptotic analysis of training overparameterized neural networks; to this literature, we contribute essentially optimal rates of concentration for the neural tangent kernel of deep fully-connected networks, requiring width $n \gtrsim L\,\mathrm{poly}(d_0)$ to achieve uniform concentration of the initial kernel over a $d_0$-dimensional submanifold of the unit sphere $\mathbb{S}^{n_0-1}$, and a nonasymptotic framework for establishing generalization of networks trained in the NTK regime with structured data. The proof makes heavy use of martingale concentration to optimally treat statistical dependencies across layers of the initial random network. This approach should be of use in establishing similar results for other network architectures.