Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks

Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks
复制标题

DOI:
--
复制
发表时间:
2019-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Mingchen Li;M. Soltanolkotabi;Samet Oymak
Mingchen Li;M. Soltanolkotabi;Samet Oymak
中科院分区:
其他
文献类型:
--
作者:
Mingchen Li;M. Soltanolkotabi;Samet Oymak

文献摘要

被引文献

相似文献

现代神经网络通常是在过参数化机制中训练的,其中模型的参数远远超过训练数据的大小。原则上,这种神经网络具有(过)拟合任何标签集(包括纯噪声)的能力。尽管如此,有些矛盾的是,通过一阶方法训练的神经网络模型仍然能够很好地预测尚未见过的测试数据。本文朝着揭开这一现象的神秘面纱迈出了一步。在一个丰富的数据集模型下,我们证明了梯度下降对恒定分数的标签上的噪声/腐败具有可证明的鲁棒性,尽管过度参数化。特别是,我们证明:(i)在最初的几次迭代中,更新仍然在初始化附近,梯度下降只适合正确的标签,基本上忽略了嘈杂的标签。(ii)为了开始过拟合到噪声标签,网络必须偏离初始化相当远,这只能在更多的迭代之后发生。总之,这些结果表明,具有早期停止的梯度下降可证明对标签噪声具有鲁棒性,并揭示了深度网络的经验鲁棒性以及通常采用的防止过拟合的算法。
Modern neural networks are typically trained in an over-parameterized regime where the parameters of the model far exceed the size of the training data. Such neural networks in principle have the capacity to (over)fit any set of labels including pure noise. Despite this, somewhat paradoxically, neural network models trained via first-order methods continue to predict well on yet unseen test data. This paper takes a step towards demystifying this phenomena. Under a rich dataset model, we show that gradient descent is provably robust to noise/corruption on a constant fraction of the labels despite overparameterization. In particular, we prove that: (i) In the first few iterations where the updates are still in the vicinity of the initialization gradient descent only fits to the correct labels essentially ignoring the noisy labels. (ii) to start to overfit to the noisy labels network must stray rather far from from the initialization which can only occur after many more iterations. Together, these results show that gradient descent with early stopping is provably robust to label noise and shed light on the empirical robustness of deep networks as well as commonly adopted heuristics to prevent overfitting.