Partial differential equations for training deep neural networks
Partial differential equations for training deep neural networks
复制标题
用于训练深度神经网络的偏微分方程
DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
G. Carlier
中科院分区:
文献类型:
--
作者:
P. Chaudhari;Adam M. Oberman;S. Osher;Stefano Soatto;G. Carlier
This paper establishes a connection between non-convex optimization and nonlinear partial differential equations (PDEs). We interpret empirically successful relaxation techniques motivated from statistical physics for training deep neural networks as solutions of a viscous Hamilton-Jacobi (HJ) PDE. The underlying stochastic control interpretation allows us to prove that these techniques perform better than stochastic gradient descent. Our analysis provides insight into the geometry of the energy landscape and suggests new algorithms based on the non-viscous Hamilton-Jacobi PDE that can effectively tackle the high dimensionality of modern neural networks.