Reconciling modern machine-learning practice and the classical bias-variance trade-off

Reconciling modern machine-learning practice and the classical bias-variance trade-off
复制标题

DOI:
10.1073/pnas.1903070116
复制
发表时间:
2019-08-06
影响因子:
11.1
通讯作者:
Mandal, Soumik
Mandal, Soumik
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Belkin, Mikhail;Hsu, Daniel;Mandal, Soumik

文献摘要

被引文献

相似文献

机器学习的突破性进展正在迅速改变科学和社会,但我们对这项技术的基本理解却远远落后。事实上,该领域的核心原则之一,偏差-方差权衡,似乎与现代机器学习实践中使用的方法的观察行为不一致。偏差-方差权衡意味着模型应该平衡欠拟合和过拟合:足够丰富以表达数据中的底层结构,并且足够简单以避免拟合虚假模式。然而,在现代实践中,诸如神经网络的非常丰富的模型被训练为精确拟合(即,内插)数据。传统上,这样的模型会被认为是过拟合的,但它们通常在测试数据上获得很高的精度。这种明显的矛盾引发了人们对机器学习的数学基础及其与从业者的相关性的质疑。在本文中,我们调和的经典理解和现代实践中的一个统一的性能曲线。这条“双下降”曲线包含了教科书中的U形偏差-方差权衡曲线,显示了增加模型容量超过插值点如何提高性能。我们为广泛的模型和数据集提供了双下降存在和普遍存在的证据,并为它的出现提供了一种机制。机器学习模型的性能和结构之间的这种联系揭示了经典分析的局限性,并对机器学习的理论和实践都有影响。
Breakthroughs in machine learning are rapidly changing science and society, yet our fundamental understanding of this technology has lagged far behind. Indeed, one of the central tenets of the field, the bias-variance trade-off, appears to be at odds with the observed behavior of methods used in modern machine-learning practice. The bias-variance trade-off implies that a model should balance underfitting and overfitting: Rich enough to express underlying structure in data and simple enough to avoid fitting spurious patterns. However, in modern practice, very rich models such as neural networks are trained to exactly fit (i.e., interpolate) the data. Classically, such models would be considered overfitted, and yet they often obtain high accuracy on test data. This apparent contradiction has raised questions about the mathematical foundations of machine learning and their relevance to practitioners. In this paper, we reconcile the classical understanding and the modern practice within a unified performance curve. This "double-descent" curve subsumes the textbook U-shaped bias-variance trade-off curve by showing how increasing model capacity beyond the point of interpolation results in improved performance. We provide evidence for the existence and ubiquity of double descent for a wide spectrum of models and datasets, and we posit a mechanism for its emergence. This connection between the performance and the structure of machine-learning models delineates the limits of classical analyses and has implications for both the theory and the practice of machine learning.