Partial differential equations for training deep neural networks

Partial differential equations for training deep neural networks
复制标题

用于训练深度神经网络的偏微分方程

DOI:
--
复制
发表时间:
2017
期刊:
Asilomar Conference on Signals, Systems and Computers
影响因子:
--
通讯作者:
G. Carlier
G. Carlier
中科院分区:
--
文献类型:
--
作者:
P. Chaudhari;Adam M. Oberman;S. Osher;Stefano Soatto;G. Carlier

文献摘要

被引文献

相似文献

建立了非凸优化与非线性偏微分方程组之间的联系。我们将经验上成功的松弛技术解释为粘性的Hamilton-Jacobi(HJ)偏微分方程解,该松弛技术源于统计物理学,用于训练深度神经网络。潜在的随机控制解释使我们能够证明这些技术比随机梯度下降更好地执行。我们的分析提供了对能量格局几何结构的洞察,并提出了基于非粘性Hamilton-Jacobi偏微分方程的新算法,该算法可以有效地解决现代神经网络的高维问题。
This paper establishes a connection between non-convex optimization and nonlinear partial differential equations (PDEs). We interpret empirically successful relaxation techniques motivated from statistical physics for training deep neural networks as solutions of a viscous Hamilton-Jacobi (HJ) PDE. The underlying stochastic control interpretation allows us to prove that these techniques perform better than stochastic gradient descent. Our analysis provides insight into the geometry of the energy landscape and suggests new algorithms based on the non-viscous Hamilton-Jacobi PDE that can effectively tackle the high dimensionality of modern neural networks.