Loss landscapes and optimization in over-parameterized non-linear systems and neural networks

Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
复制标题

DOI:
10.1016/j.acha.2021.12.009
复制
发表时间:
2022-04-25
影响因子:
2.5
通讯作者:
Belkin, Mikhail
Belkin, Mikhail
中科院分区:
数学1区
文献类型:
--
作者:
Liu, Chaoyue;Zhu, Libin;Belkin, Mikhail

文献摘要

被引文献

相似文献

深度学习的成功在很大程度上要归功于基于梯度的优化方法应用于大型神经网络的显著效果。这项工作的目的是提出一个现代的观点和一个通用的数学框架,用于过度参数化的机器学习模型和非线性方程系统中的损失景观和有效优化,这种设置包括过度参数化的深度神经网络。我们最初的观察是,对应于这样的系统的优化景观通常不是凸的,即使是在全局最小值附近,我们称之为本质非凸性的条件。相反,我们认为它们满足PL*,这是Polyak-Lojasiewicz条件[32,25]在大多数(但不是全部)参数空间上的变体,它保证了解的存在性和(随机)梯度下降(SGD/GD)的有效优化。这些系统的PL* 条件与非线性系统相关的切线核的条件数密切相关,表明基于PL* 的非线性理论如何与过参数化线性方程的经典分析并行。我们证明了宽神经网络满足PL* 条件,这解释了(S)GD收敛到全局最小值。最后,我们提出了一个放松的PL* 条件适用于“几乎”过参数化系统。(C)2021爱思唯尔公司All rights reserved.
The success of deep learning is due, to a large extent, to the remarkable effectiveness of gradient-based optimization methods applied to large neural networks. The purpose of this work is to propose a modern view and a general mathematical framework for loss landscapes and efficient optimization in over-parameterized machine learning models and systems of non-linear equations, a setting that includes over-parameterized deep neural networks. Our starting observation is that optimization landscapes corresponding to such systems are generally not convex, even locally around a global minimum, a condition we call essential non-convexity. We argue that instead they satisfy PL*, a variant of the Polyak-Lojasiewicz condition [32,25] on most (but not all) of the parameter space, which guarantees both the existence of solutions and efficient optimization by (stochastic) gradient descent (SGD/GD). The PL* condition of these systems is closely related to the condition number of the tangent kernel associated to a non-linear system showing how a PL*-based non-linear theory parallels classical analyses of over-parameterized linear equations. We show that wide neural networks satisfy the PL* condition, which explains the (S)GD convergence to a global minimum. Finally we propose a relaxation of the PL* condition applicable to "almost " over-parameterized systems. (C)& nbsp;2021 Elsevier Inc. All rights reserved.