The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo

The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo
复制标题

DOI:
10.5555/2627435.2638586
复制
发表时间:
2011-11
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
M. Hoffman;A. Gelman
M. Hoffman;A. Gelman
中科院分区:
其他
文献类型:
--
作者:
M. Hoffman;A. Gelman

文献摘要

被引文献

相似文献

Hamilton Monte Carlo(HMC)是一种马尔可夫链蒙特卡罗(MCMC)算法,通过采取一系列由一阶梯度信息通知的步骤,避免了困扰许多MCMC方法的随机游走行为和对相关参数的敏感性。这些特征使它能够比随机游走大都会或吉布斯抽样等简单方法更快地收敛到高维目标分布。然而,HMC的性能对两个用户指定的参数高度敏感:步长和所需的步长数L。特别地,如果L太小,则算法表现出不期望的随机游走行为,而如果L太大,则算法浪费计算。我们引入了无U形转弯采样器(NUTS),这是HMC的一个扩展,它消除了设置步长L的需要。NUTS使用递归算法来构建一组可能的候选点,这些候选点跨越目标分布的宽范围,当它开始往回走并回溯其步骤时自动停止。经验上,NUTS执行至少与良好调优的标准HMC方法一样有效,有时更有效,而不需要用户干预或昂贵的调优运行。我们还推导出一种基于原始-对偶平均的动态调整步长参数的方法。因此,NUTS可以在根本不需要手动调谐的情况下使用。NUTS还适用于需要高效“交钥匙”采样算法的BUGS式自动推理引擎等应用。
Hamiltonian Monte Carlo (HMC) is a Markov chain Monte Carlo (MCMC) algorithm that avoids the random walk behavior and sensitivity to correlated parameters that plague many MCMC methods by taking a series of steps informed by first-order gradient information. These features allow it to converge to high-dimensional target distributions much more quickly than simpler methods such as random walk Metropolis or Gibbs sampling. However, HMC's performance is highly sensitive to two user-specified parameters: a step size {\epsilon} and a desired number of steps L. In particular, if L is too small then the algorithm exhibits undesirable random walk behavior, while if L is too large the algorithm wastes computation. We introduce the No-U-Turn Sampler (NUTS), an extension to HMC that eliminates the need to set a number of steps L. NUTS uses a recursive algorithm to build a set of likely candidate points that spans a wide swath of the target distribution, stopping automatically when it starts to double back and retrace its steps. Empirically, NUTS perform at least as efficiently as and sometimes more efficiently than a well tuned standard HMC method, without requiring user intervention or costly tuning runs. We also derive a method for adapting the step size parameter {\epsilon} on the fly based on primal-dual averaging. NUTS can thus be used with no hand-tuning at all. NUTS is also suitable for applications such as BUGS-style automatic inference engines that require efficient "turnkey" sampling algorithms.