Control variates for stochastic gradient MCMC

Control variates for stochastic gradient MCMC
复制标题

DOI:
10.1007/s11222-018-9826-2
复制
发表时间:
2017-06
影响因子:
2.2
通讯作者:
Jack Baker;P. Fearnhead;E. Fox;C. Nemeth
Jack Baker;P. Fearnhead;E. Fox;C. Nemeth
中科院分区:
数学2区
文献类型:
--
作者:
Jack Baker;P. Fearnhead;E. Fox;C. Nemeth

文献摘要

相似文献

众所周知,马尔可夫链蒙特卡罗 (MCMC) 方法随数据集大小的扩展性很差。解决此问题的一类流行方法是随机梯度 MCMC (SGMCMC)。这些方法使用对数后验梯度的噪声估计,这减少了算法的每次迭代计算成本。 Despite this, there are a number of results suggesting that stochastic gradient Langevin dynamics (SGLD), probably the most popular of these methods, still has computational cost proportional to the dataset size.我们建议随机梯度 MCMC 的替代对数后验梯度估计,它使用控制变量来减少方差。 We analyse SGLD using this gradient estimate, and show that, under log-concavity assumptions on the target distribution, the computational cost required for a given level of accuracy is independent of the dataset size.接下来,我们展示了一种不同的控制变量技术,称为零方差控制变量,可以免费应用于 SGMCMC 算法。此后处理步骤通过减少 MCMC 输出的方差来改进算法的推理。零方差控制变量依赖于对数后验的梯度;我们通过将其替换为 SGMCMC 计算的噪声梯度估计来探索方差减少的影响。
It is well known that Markov chain Monte Carlo (MCMC) methods scale poorly with dataset size. A popular class of methods for solving this issue is stochastic gradient MCMC (SGMCMC). These methods use a noisy estimate of the gradient of the log-posterior, which reduces the per iteration computational cost of the algorithm. Despite this, there are a number of results suggesting that stochastic gradient Langevin dynamics (SGLD), probably the most popular of these methods, still has computational cost proportional to the dataset size. We suggest an alternative log-posterior gradient estimate for stochastic gradient MCMC which uses control variates to reduce the variance. We analyse SGLD using this gradient estimate, and show that, under log-concavity assumptions on the target distribution, the computational cost required for a given level of accuracy is independent of the dataset size. Next, we show that a different control-variate technique, known as zero variance control variates, can be applied to SGMCMC algorithms for free. This postprocessing step improves the inference of the algorithm by reducing the variance of the MCMC output. Zero variance control variates rely on the gradient of the log-posterior; we explore how the variance reduction is affected by replacing this with the noisy gradient estimate calculated by SGMCMC.