Learning One-hidden-layer ReLU Networks via Gradient Descent

Learning One-hidden-layer ReLU Networks via Gradient Descent
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
--
影响因子:
--
通讯作者:
Xiao Zhang;Yaodong Yu;Lingxiao Wang;Quanquan Gu
Xiao Zhang;Yaodong Yu;Lingxiao Wang;Quanquan Gu
中科院分区:
其他
文献类型:
--
作者:
Xiao Zhang;Yaodong Yu;Lingxiao Wang;Quanquan Gu

文献摘要

被引文献

相似文献

我们研究具有整流线性单元(RELU)激活函数的一层隐形神经网络的问题,其中输入是从标准高斯分布中采样的,并且输出是从嘈杂的教师网络中产生的。我们分析了基于经验风险最小化训练这种神经网络的梯度下降的性能,并提供依赖算法的保证。特别是,我们证明张量初始化后下降可以以线性速率收敛到地面真相参数,直到一定的统计误差。据我们所知,这是表征具有多个神经元的一层层次恢复网络的恢复保证的第一项工作。数值实验验证了我们的理论发现。
We study the problem of learning one-hidden-layer neural networks with Rectified Linear Unit (ReLU) activation function, where the inputs are sampled from standard Gaussian distribution and the outputs are generated from a noisy teacher network. We analyze the performance of gradient descent for training such kind of neural networks based on empirical risk minimization, and provide algorithm-dependent guarantees. In particular, we prove that tensor initialization followed by gradient descent can converge to the ground-truth parameters at a linear rate up to some statistical error. To the best of our knowledge, this is the first work characterizing the recovery guarantee for practical learning of one-hidden-layer ReLU networks with multiple neurons. Numerical experiments verify our theoretical findings.