Breaking Reversibility Accelerates Langevin Dynamics for Non-Convex Optimization

Breaking Reversibility Accelerates Langevin Dynamics for Non-Convex Optimization
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Xuefeng Gao;M. Gürbüzbalaban;Lingjiong Zhu
Xuefeng Gao;M. Gürbüzbalaban;Lingjiong Zhu
中科院分区:
其他
文献类型:
--
作者:
Xuefeng Gao;M. Gürbüzbalaban;Lingjiong Zhu

文献摘要

被引文献

相似文献

朗之万动力学(LD)已被证明是一种强大的技术,用于优化非凸目标,作为一种有效的算法来寻找局部极小值,同时最终在更长的时间尺度上访问全局极小值。LD基于一阶Langevin扩散,其在时间上是可逆的。我们研究了两个变种,是基于不可逆的朗之万扩散:欠阻尼朗之万动力学(ULD)和朗之万动力学与非对称漂移(NLD)。采用Tzen等人的技术。(2018)对于LD到不可逆扩散,我们表明,对于给定的局部最小值,在距离初始化任意距离内,具有很高的概率,ULD轨迹在递归时间内结束于该局部最小值的小邻域之外的某处,该递归时间取决于Hessian在局部最小值处的最小特征值,或者它们进入该局部最小值。并在那里停留一个潜在的指数级长的逃逸时间。ULD算法改进了Tzen等人(2018)在局部最小值处对Hessian最小特征值的依赖性方面为LD获得的递归时间。NLD算法也得到了类似的结果和改进。我们还表明,不可逆的变体可以退出吸引盆的局部最小值更快地在离散时间时,目标有两个局部极小值分离的鞍点和量化的改善量。我们的分析表明,不可逆的朗之万算法是
Langevin dynamics (LD) has been proven to be a powerful technique for optimizing a non-convex objective as an efficient algorithm to find local minima while eventually visiting a global minimum on longer time-scales. LD is based on the first-order Langevin diffusion which is reversible in time. We study two variants that are based on non-reversible Langevin diffusions: the underdamped Langevin dynamics (ULD) and the Langevin dynamics with a non-symmetric drift (NLD). Adopting the techniques of Tzen et al. (2018) for LD to non-reversible diffusions, we show that for a given local minimum that is within an arbitrary distance from the initialization, with high probability, either the ULD trajectory ends up somewhere outside a small neighborhood of this local minimum within a recurrence time which depends on the smallest eigenvalue of the Hessian at the local minimum or they enter this neighborhood by the recurrence time and stay there for a potentially exponentially long escape time. The ULD algorithm improves upon the recurrence time obtained for LD in Tzen et al. (2018) with respect to the dependency on the smallest eigenvalue of the Hessian at the local minimum. Similar results and improvements are obtained for the NLD algorithm. We also show that non-reversible variants can exit the basin of attraction of a local minimum faster in discrete time when the objective has two local minima separated by a saddle point and quantify the amount of improvement. Our analysis suggests that non-reversible Langevin algorithms are