Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting

Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
复制标题

DOI:
10.48550/arxiv.2207.06569
复制
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Neil Rohit Mallinar;James B. Simon;Amirhesam Abedsoltan;Parthe Pandit;M. Belkin;Preetum Nakkiran
Neil Rohit Mallinar;James B. Simon;Amirhesam Abedsoltan;Parthe Pandit;M. Belkin;Preetum Nakkiran
中科院分区:
其他
文献类型:
--
作者:
Neil Rohit Mallinar;James B. Simon;Amirhesam Abedsoltan;Parthe Pandit;M. Belkin;Preetum Nakkiran

文献摘要

被引文献

相似文献

超参数化神经网络的实际成功激发了最近对插值方法的科学研究,这些方法完全符合其训练数据。某些插值方法,包括神经网络,可以拟合有噪声的训练数据,而不会有灾难性的糟糕测试性能,这与统计学习理论的标准直觉背道而驰。为了解释这一点,最近的一系列工作研究了良性过拟合,这是一种即使在存在噪声的情况下,一些插值方法也接近贝叶斯最优的现象。在这项工作中,我们认为,虽然良性过拟合一直是有益的和富有成效的研究,许多真实的插值方法,如神经网络不适合良性:适度的噪声在训练集导致非零(但非无限)的过度风险在测试时,这意味着这些模型既不是良性的,也不是灾难性的,而是在一个中间政权下降。我们称这种中间状态为回火过拟合,并对其进行了系统的研究。我们首先探讨这一现象的背景下,核(岭)回归(KR)得到的条件下,KR表现出的三种行为的岭参数和核特征谱。我们发现具有幂律谱的核,包括拉普拉斯核和ReLU神经正切核,表现出回火过拟合。然后,我们通过分类学的透镜对深度神经网络进行了实证研究,发现那些经过插值训练的神经网络是温和的,而那些提前停止的神经网络是良性的。我们希望我们的工作能让我们对现代学习中的过拟合有更深入的理解。
The practical success of overparameterized neural networks has motivated the recent scientific study of interpolating methods, which perfectly fit their training data. Certain interpolating methods, including neural networks, can fit noisy training data without catastrophically bad test performance, in defiance of standard intuitions from statistical learning theory. Aiming to explain this, a body of recent work has studied benign overfitting, a phenomenon where some interpolating methods approach Bayes optimality, even in the presence of noise. In this work we argue that while benign overfitting has been instructive and fruitful to study, many real interpolating methods like neural networks do not fit benignly: modest noise in the training set causes nonzero (but non-infinite) excess risk at test time, implying these models are neither benign nor catastrophic but rather fall in an intermediate regime. We call this intermediate regime tempered overfitting, and we initiate its systematic study. We first explore this phenomenon in the context of kernel (ridge) regression (KR) by obtaining conditions on the ridge parameter and kernel eigenspectrum under which KR exhibits each of the three behaviors. We find that kernels with powerlaw spectra, including Laplace kernels and ReLU neural tangent kernels, exhibit tempered overfitting. We then empirically study deep neural networks through the lens of our taxonomy, and find that those trained to interpolation are tempered, while those stopped early are benign. We hope our work leads to a more refined understanding of overfitting in modern learning.