Optimal Regularization Can Mitigate Double Descent

Optimal Regularization Can Mitigate Double Descent
复制标题

DOI:
--
复制
发表时间:
2020-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Preetum Nakkiran;Prayaag Venkat;S. Kakade;Tengyu Ma
Preetum Nakkiran;Prayaag Venkat;S. Kakade;Tengyu Ma
中科院分区:
其他
文献类型:
--
作者:
Preetum Nakkiran;Prayaag Venkat;S. Kakade;Tengyu Ma

文献摘要

被引文献

相似文献

最近的实证和理论研究表明,许多学习算法——从线性回归到神经网络——在样本大小和模型大小等数量上都可以具有非单调的测试性能。这种惊人的现象通常被称为“双重下降”,它提出了一个问题,即我们是否需要重新思考我们目前对泛化的理解。在本工作中,我们研究了是否可以使用最优正则化来避免双下降现象。从理论上证明了对于某些具有各向同性数据分布的线性回归模型,随着样本量或模型大小的增加,最优调谐$\ell_2$正则化达到单调的检验性能。我们也从经验上证明了最优调整的$\ell_2$正则化可以缓解更一般的模型(包括神经网络)的双重下降。我们的结果表明,在适当调整正则化的背景下,研究各种算法的测试风险缩放也可能是有益的。
Recent empirical and theoretical studies have shown that many learning algorithms -- from linear regression to neural networks -- can have test performance that is non-monotonic in quantities such the sample size and model size. This striking phenomenon, often referred to as "double descent", has raised questions of if we need to re-think our current understanding of generalization. In this work, we study whether the double-descent phenomenon can be avoided by using optimal regularization. Theoretically, we prove that for certain linear regression models with isotropic data distribution, optimally-tuned $\ell_2$ regularization achieves monotonic test performance as we grow either the sample size or the model size. We also demonstrate empirically that optimally-tuned $\ell_2$ regularization can mitigate double descent for more general models, including neural networks. Our results suggest that it may also be informative to study the test risk scalings of various algorithms in the context of appropriately tuned regularization.