Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data

Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Spencer Frei;Niladri S. Chatterji;P. Bartlett
Spencer Frei;Niladri S. Chatterji;P. Bartlett
中科院分区:
其他
文献类型:
--
作者:
Spencer Frei;Niladri S. Chatterji;P. Bartlett

文献摘要

被引文献

相似文献

良性过拟合,即插值模型在存在噪声数据的情况下泛化良好的现象,首先在梯度下降训练的神经网络模型中观察到。为了更好地理解这种经验观察,我们考虑了随机初始化后逻辑损失的梯度下降插值训练的两层神经网络的泛化误差。我们假设数据来自分离良好的类条件对数凹分布,并允许训练标签的恒定部分被对手破坏。我们表明,在这种情况下,神经网络表现出良性的过拟合:它们可以被驱动到零训练误差,完美地拟合任何有噪声的训练标签,同时实现最小最大最优测试误差。与之前需要线性或基于内核的预测器的良性过拟合工作相比,我们的分析在模型和学习动态基本上都是非线性的情况下成立。
Benign overfitting, the phenomenon where interpolating models generalize well in the presence of noisy data, was first observed in neural network models trained with gradient descent. To better understand this empirical observation, we consider the generalization error of two-layer neural networks trained to interpolation by gradient descent on the logistic loss following random initialization. We assume the data comes from well-separated class-conditional log-concave distributions and allow for a constant fraction of the training labels to be corrupted by an adversary. We show that in this setting, neural networks exhibit benign overfitting: they can be driven to zero training error, perfectly fitting any noisy training labels, and simultaneously achieve minimax optimal test error. In contrast to previous work on benign overfitting that require linear or kernel-based predictors, our analysis holds in a setting where both the model and learning dynamics are fundamentally nonlinear.