Stochastic Gradient Descent as Approximate Bayesian Inference

Stochastic Gradient Descent as Approximate Bayesian Inference
复制标题

DOI:
--
复制
发表时间:
2017-04
期刊:
ArXiv
影响因子:
--
通讯作者:
S. Mandt;M. Hoffman;D. Blei
S. Mandt;M. Hoffman;D. Blei
中科院分区:
其他
文献类型:
--
作者:
S. Mandt;M. Hoffman;D. Blei

文献摘要

被引文献

相似文献

具有恒定学习率(恒定SGD)的随机梯度下降模拟具有平稳分布的马尔可夫链。从这个角度来看,我们得到了一些新的结果。(1)我们表明,常数SGD可以作为一个近似的贝叶斯后验推理算法。具体来说,我们展示了如何调整常数SGD的调谐参数,以最好地匹配平稳分布的后验,最大限度地减少这两个分布之间的Kullback-Leibler分歧。(2)我们证明了恒定SGD产生了一个新的变分EM算法,优化超参数在复杂的概率模型。(3)我们还提出了SGD与动量采样,并显示如何相应地调整阻尼系数。(4)我们分析MCMC算法。对于Langevin动力学和随机梯度Fisher评分,我们量化了由于有限学习率而导致的近似误差。最后(5),我们使用随机过程的观点来给出为什么Polyak平均是最优的简短证明。基于这一思想,我们提出了一个可扩展的近似MCMC算法,平均随机梯度采样。
Stochastic Gradient Descent with a constant learning rate (constant SGD) simulates a Markov chain with a stationary distribution. With this perspective, we derive several new results. (1) We show that constant SGD can be used as an approximate Bayesian posterior inference algorithm. Specifically, we show how to adjust the tuning parameters of constant SGD to best match the stationary distribution to a posterior, minimizing the Kullback-Leibler divergence between these two distributions. (2) We demonstrate that constant SGD gives rise to a new variational EM algorithm that optimizes hyperparameters in complex probabilistic models. (3) We also propose SGD with momentum for sampling and show how to adjust the damping coefficient accordingly. (4) We analyze MCMC algorithms. For Langevin Dynamics and Stochastic Gradient Fisher Scoring, we quantify the approximation errors due to finite learning rates. Finally (5), we use the stochastic process perspective to give a short proof of why Polyak averaging is optimal. Based on this idea, we propose a scalable approximate MCMC algorithm, the Averaged Stochastic Gradient Sampler.