Guaranteed Recovery of One-Hidden-Layer Neural Networks via Cross Entropy

Guaranteed Recovery of One-Hidden-Layer Neural Networks via Cross Entropy
复制标题

DOI:
10.1109/tsp.2020.2993153
复制
发表时间:
2018-02
影响因子:
5.4
通讯作者:
H. Fu;Yuejie Chi;Yingbin Liang
H. Fu;Yuejie Chi;Yingbin Liang
中科院分区:
工程技术1区
文献类型:
--
作者:
H. Fu;Yuejie Chi;Yingbin Liang

文献摘要

相似文献

我们研究了数据分类的模型恢复,其中训练标签是从具有sigmoid激活的一个隐藏层神经网络(也称为单层前馈网络)中生成的,目标是恢复神经网络的权重。我们考虑两种网络模型,全连接网络(FCN)和非重叠卷积神经网络(CNN)。我们证明了高斯输入,基于交叉熵的经验风险表现出强的凸性和光滑性,一致地在一个局部邻域的地面真理,只要样本的复杂性是足够大的。这意味着如果在这个邻域中初始化,梯度下降线性收敛到一个临界点,该临界点可证明接近地面真理。此外,我们表明这样的初始化可以通过张量方法获得。这为经验风险最小化建立了全局收敛保证,使用交叉熵通过梯度下降学习一个隐藏层神经网络,在接近最佳的样本和计算复杂度相对于网络输入维度,没有不切实际的假设,如需要一组新的样本在每次迭代。
We study model recovery for data classification, where the training labels are generated from a one-hidden-layer neural network with sigmoid activations, also known as a single-layer feedforward network, and the goal is to recover the weights of the neural network. We consider two network models, the fully-connected network (FCN) and the non-overlapping convolutional neural network (CNN). We prove that with Gaussian inputs, the empirical risk based on cross entropy exhibits strong convexity and smoothness uniformly in a local neighborhood of the ground truth, as soon as the sample complexity is sufficiently large. This implies that if initialized in this neighborhood, gradient descent converges linearly to a critical point that is provably close to the ground truth. Furthermore, we show such an initialization can be obtained via the tensor method. This establishes the global convergence guarantee for empirical risk minimization using cross entropy via gradient descent for learning one-hidden-layer neural networks, at the near-optimal sample and computational complexity with respect to the network input dimension without unrealistic assumptions such as requiring a fresh set of samples at each iteration.