Langevin Dynamics with Continuous Tempering for High-dimensional Non-convex Optimization

Langevin Dynamics with Continuous Tempering for High-dimensional Non-convex Optimization
复制标题

Langevin Dynamics 与连续回火的高维非凸优化

DOI:
10.17863/cam.25294
复制
发表时间:
2017
期刊:
ArXiv
影响因子:
--
通讯作者:
Rafał K. Mantiuk
Rafał K. Mantiuk
中科院分区:
--
文献类型:
--
作者:
Nanyang Ye;Zhanxing Zhu;Rafał K. Mantiuk

文献摘要

参考文献

被引文献

相似文献

最小化非凸和高维目标函数具有挑战性,特别是在训练现代深度神经网络时。本文提出了一种新的方法,将训练过程分为两个连续的阶段,以获得更好的泛化性能:贝叶斯采样和随机优化。第一阶段是探索能源格局并捕获“胖”模式;第二阶段是微调从第一阶段学习的参数。在贝叶斯学习阶段,我们将连续回火和随机逼近应用到朗之万动力学中,以创建一个高效且有效的采样器,其中温度根据设计的“温度动力学”自动调整。正如我们的理论分析和实证实验所示,这些策略可以克服早期陷入不良局部极小值的挑战,并在各种类型的神经网络中取得了显着的改进。
Minimizing non-convex and high-dimensional objective functions are challenging, especially when training modern deep neural networks. In this paper, a novel approach is proposed which divides the training process into two consecutive phases to obtain better generalization performance: Bayesian sampling and stochastic optimization. The first phase is to explore the energy landscape and to capture the "fat" modes; and the second one is to fine-tune the parameter learned from the first phase. In the Bayesian learning phase, we apply continuous tempering and stochastic approximation into the Langevin dynamics to create an efficient and effective sampler, in which the temperature is adjusted automatically according to the designed "temperature dynamics". These strategies can overcome the challenge of early trapping into bad local minima and have achieved remarkable improvements in various types of neural networks as shown in our theoretical analysis and empirical experiments.
DOI: 10.1137/090770527
发表时间: 2009-08
期刊: SIAM J. Numer. Anal.
影响因子: --
作者:
Jonathan C. Mattingly;A. Stuart;M. Tretyakov
通讯作者: Jonathan C. Mattingly;A. Stuart;M. Tretyakov
DOI: 10.1103/physreve.91.061301
发表时间: 2015-06-19
期刊: PHYSICAL REVIEW E
影响因子: 2.4
作者:
Gobbo, Gianpaolo;Leimkuhler, Benedict J.
通讯作者: Leimkuhler, Benedict J.