Proxy Convexity: A Unified Framework for the Analysis of Neural Networks Trained by Gradient Descent

Proxy Convexity: A Unified Framework for the Analysis of Neural Networks Trained by Gradient Descent
复制标题

DOI:
--
复制
发表时间:
2021-06
影响因子:
20.6
通讯作者:
Spencer Frei;Quanquan Gu
Spencer Frei;Quanquan Gu
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Spencer Frei;Quanquan Gu

文献摘要

被引文献

相似文献

虽然学习神经网络的优化目标是高度非凸的,但基于梯度的方法在学习神经网络的实践中已经取得了巨大的成功。这种并列关系导致了最近一些关于通过梯度下降训练的神经网络的可证明保证的研究。不幸的是,这些作品中的技术往往高度特定于每个问题中的特定设置,因此很难在不同的设置中进行推广。针对文献中的这一缺陷,我们提出了一个用于神经网络训练分析的统一的非凸优化框架。我们引入了代理凸性和代理Polyak-Lojasiewicz(PL)不等式的概念,如果原始目标函数诱导出一个在使用梯度方法时隐式最小化的代理目标函数,则这两个概念是满足的。我们证明了满足代理凸性或代理PL不等式的目标的梯度下降导致了对代理目标函数的有效保证。我们进一步证明了由梯度下降训练的神经网络的许多现有的保证可以通过代理凸性和代理PL不等式来统一。
Although the optimization objectives for learning neural networks are highly non-convex, gradient-based methods have been wildly successful at learning neural networks in practice. This juxtaposition has led to a number of recent studies on provable guarantees for neural networks trained by gradient descent. Unfortunately, the techniques in these works are often highly specific to the particular setup in each problem, making it difficult to generalize across different settings. To address this drawback in the literature, we propose a unified non-convex optimization framework for the analysis of neural network training. We introduce the notions of proxy convexity and proxy Polyak-Lojasiewicz (PL) inequalities, which are satisfied if the original objective function induces a proxy objective function that is implicitly minimized when using gradient methods. We show that gradient descent on objectives satisfying proxy convexity or the proxy PL inequality leads to efficient guarantees for proxy objective functions. We further show that many existing guarantees for neural networks trained by gradient descent can be unified through proxy convexity and proxy PL inequalities.