Deep Networks Provably Classify Data on Curves

Deep Networks Provably Classify Data on Curves
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
--
影响因子:
--
通讯作者:
Tingran Wang;Sam Buchanan;D. Gilboa;John N. Wright
Tingran Wang;Sam Buchanan;D. Gilboa;John N. Wright
中科院分区:
其他
文献类型:
--
作者:
Tingran Wang;Sam Buchanan;D. Gilboa;John N. Wright

文献摘要

相似文献

具有低维非线性结构的数据在工程和科学问题中普遍存在。我们研究了这样的结构的模型问题-一个二进制分类任务,使用一个深的全连接神经网络分类数据从两个不相交的光滑曲线上的单位球。除了温和的正则性条件,我们没有限制的配置曲线。我们证明,当(i)网络深度相对于确定问题难度的某些几何属性较大,以及(ii)网络宽度和样本数量在深度上是多项式时,随机初始化梯度下降快速学习以高概率正确分类两条曲线上的所有点。据我们所知,这是深度网络的第一个泛化保证,其非线性数据仅取决于内在数据属性。我们的分析通过减少神经正切核(NTK)机制中的动态来进行,其中网络深度在解决分类问题中扮演着拟合资源的角色。特别是,通过对NTK的衰减特性进行细粒度控制,我们证明了当网络足够深时,NTK可以在流形上局部近似于一个几何不变的算子,并在光滑函数上稳定地反演,这保证了收敛和推广。
Data with low-dimensional nonlinear structure are ubiquitous in engineering and scientific problems. We study a model problem with such structure—a binary classification task that uses a deep fully-connected neural network to classify data drawn from two disjoint smooth curves on the unit sphere. Aside from mild regularity conditions, we place no restrictions on the configuration of the curves. We prove that when (i) the network depth is large relative to certain geometric properties that set the difficulty of the problem and (ii) the network width and number of samples are polynomial in the depth, randomly-initialized gradient descent quickly learns to correctly classify all points on the two curves with high probability. To our knowledge, this is the first generalization guarantee for deep networks with nonlinear data that depends only on intrinsic data properties. Our analysis proceeds by a reduction to dynamics in the neural tangent kernel (NTK) regime, where the network depth plays the role of a fitting resource in solving the classification problem. In particular, via fine-grained control of the decay properties of the NTK, we demonstrate that when the network is sufficiently deep, the NTK can be locally approximated by a translationally invariant operator on the manifolds and stably inverted over smooth functions, which guarantees convergence and generalization.