Identifying and attacking the saddle point problem in high-dimensional non-convex optimization

Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
复制标题

DOI:
--
复制
发表时间:
2014-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yann Dauphin;Razvan Pascanu;Çaglar Gülçehre;Kyunghyun Cho;S. Ganguli;Yoshua Bengio
Yann Dauphin;Razvan Pascanu;Çaglar Gülçehre;Kyunghyun Cho;S. Ganguli;Yoshua Bengio
中科院分区:
其他
文献类型:
--
作者:
Yann Dauphin;Razvan Pascanu;Çaglar Gülçehre;Kyunghyun Cho;S. Ganguli;Yoshua Bengio

文献摘要

被引文献

相似文献

许多科学和工程领域面临的一个核心挑战是在连续的高维空间中最小化非凸误差函数。梯度下降法或拟牛顿法几乎被普遍用于执行此类最小化操作,并且人们通常认为,这些局部方法难以找到全局最小值的一个主要原因是,存在大量误差远高于全局最小值的局部最小值。在此,基于统计物理学、随机矩阵理论、神经网络理论的结果以及经验证据,我们认为,一个更深层次且更严重的困难源于鞍点而非局部最小值的大量出现,在具有实际应用价值的高维问题中尤其如此。此类鞍点被高误差平台所环绕,这会极大地减慢学习速度,并给人一种存在局部最小值的错觉。基于这些观点,我们提出了一种新的二阶优化方法——无鞍牛顿法,与梯度下降法和拟牛顿法不同,它能够快速逃离高维鞍点。我们将该算法应用于深度神经网络或循环神经网络的训练,并为其优越的优化性能提供了数值证据。
A central challenge to many fields of science and engineering involves minimizing non-convex error functions over continuous, high dimensional spaces. Gradient descent or quasi-Newton methods are almost ubiquitously used to perform such minimizations, and it is often thought that a main source of difficulty for these local methods to find the global minimum is the proliferation of local minima with much higher error than the global minimum. Here we argue, based on results from statistical physics, random matrix theory, neural network theory, and empirical evidence, that a deeper and more profound difficulty originates from the proliferation of saddle points, not local minima, especially in high dimensional problems of practical interest. Such saddle points are surrounded by high error plateaus that can dramatically slow down learning, and give the illusory impression of the existence of a local minimum. Motivated by these arguments, we propose a new approach to second-order optimization, the saddle-free Newton method, that can rapidly escape high dimensional saddle points, unlike gradient descent and quasi-Newton methods. We apply this algorithm to deep or recurrent neural network training, and provide numerical evidence for its superior optimization performance.