Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation

Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
复制标题

DOI:
10.1017/s0962492921000039
复制
发表时间:
2021-05-01
期刊:
影响因子:
14.2
通讯作者:
Belkin, Mikhail
Belkin, Mikhail
中科院分区:
数学1区
文献类型:
--
作者:
Belkin, Mikhail

文献摘要

被引文献

相似文献

在过去的十年中,机器学习的数学理论远远落后于深度神经网络在实际挑战中的胜利。然而,理论与实践之间的差距正逐渐开始缩小。在这篇论文中,我将试图收集一些引人注目的、仍然不完整的数学拼图,这些拼图是在理解深度学习基础的努力中出现的。两个关键的主题将是插值和它的兄弟过参数化。插值与拟合数据,甚至噪声数据,完全对应。过参数化允许插值,并提供选择合适插值模型的灵活性。正如我们将看到的,就像物理棱镜分离光线中混合的颜色一样,插值的比喻棱镜有助于在现代机器学习的复杂画面中解开泛化和优化属性。这篇文章是在这样的信念下写的,希望对这些问题的更清晰的理解将使我们更接近深度学习和机器学习的一般理论。
In the past decade the mathematical theory of machine learning has lagged far behind the triumphs of deep neural networks on practical challenges. However, the gap between theory and practice is gradually starting to close. In this paper I will attempt to assemble some pieces of the remarkable and still incomplete mathematical mosaic emerging from the efforts to understand the foundations of deep learning. The two key themes will be interpolation and its sibling over-parametrization. Interpolation corresponds to fitting data, even noisy data, exactly. Over-parametrization enables interpolation and provides flexibility to select a suitable interpolating model. As we will see, just as a physical prism separates colours mixed within a ray of light, the figurative prism of interpolation helps to disentangle generalization and optimization properties within the complex picture of modern machine learning. This article is written in the belief and hope that clearer understanding of these issues will bring us a step closer towards a general theory of deep learning and machine learning.