Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise

Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise
复制标题

DOI:
--
复制
发表时间:
2021-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Spencer Frei;Yuan Cao;Quanquan Gu
Spencer Frei;Yuan Cao;Quanquan Gu
中科院分区:
其他
文献类型:
--
作者:
Spencer Frei;Yuan Cao;Quanquan Gu

文献摘要

相似文献

我们考虑一个任意宽度的单隐藏层泄漏ReLU网络,该网络在任意初始化后通过随机梯度下降训练。我们证明了随机梯度下降(SGD)产生的神经网络,具有竞争力的分类精度的最佳半空间分布的广泛的一类分布,包括对数凹各向同性和硬边缘分布。等价地,当数据分布是线性可分的但被对抗性标签噪声破坏时,这样的网络可以泛化,尽管有过拟合的能力。我们进行的实验表明,对于某些分布,我们的泛化范围几乎是紧的。这是第一个结果表明,当数据被对抗性标签噪声破坏时,SGD训练的过参数化神经网络可以泛化。
We consider a one-hidden-layer leaky ReLU network of arbitrary width trained by stochastic gradient descent following an arbitrary initialization. We prove that stochastic gradient descent (SGD) produces neural networks that have classification accuracy competitive with that of the best halfspace over the distribution for a broad class of distributions that includes log-concave isotropic and hard margin distributions. Equivalently, such networks can generalize when the data distribution is linearly separable but corrupted with adversarial label noise, despite the capacity to overfit. We conduct experiments which suggest that for some distributions our generalization bounds are nearly tight. This is the first result that shows that overparameterized neural networks trained by SGD can generalize when the data is corrupted with adversarial label noise.