Structured Stochastic Gradient MCMC

Structured Stochastic Gradient MCMC
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Antonios Alexos;Alex Boyd;S. Mandt
Antonios Alexos;Alex Boyd;S. Mandt
中科院分区:
其他
文献类型:
--
作者:
Antonios Alexos;Alex Boyd;S. Mandt

文献摘要

相似文献

随机梯度马尔可夫链蒙特卡罗(SGMCMC)被认为是大规模模型(如贝叶斯神经网络)中贝叶斯推理的黄金标准。由于在这些模型中,从业者面临速度与准确性的权衡,变分推理(VI)通常是首选。不幸的是,VI对后验的因子分解和函数形式都做了很强的假设。在这项工作中,我们提出了一个新的非参数变分近似,使近似后验的函数形式没有假设,并允许从业者指定确切的依赖关系的算法应该尊重或打破。该方法依赖于一种新的Langevin型算法,该算法对修改后的能量函数进行操作,其中部分潜在变量在马尔可夫链的早期迭代的样本上进行平均。通过这种方式,可以以可控的方式打破统计依赖性,从而使链更快地混合。该方案可以以“dropout“方式进一步修改,从而导致更大的可扩展性。我们在CIFAR-10,SVHN和FMNIST上测试了我们的ResNet-20方案。在所有情况下,我们发现与SG-MCMC和VI相比,收敛速度和/或最终精度有所提高。
Stochastic gradient Markov Chain Monte Carlo (SGMCMC) is considered the gold standard for Bayesian inference in large-scale models, such as Bayesian neural networks. Since practitioners face speed versus accuracy tradeoffs in these models, variational inference (VI) is often the preferable option. Unfortunately, VI makes strong assumptions on both the factorization and functional form of the posterior. In this work, we propose a new non-parametric variational approximation that makes no assumptions about the approximate posterior's functional form and allows practitioners to specify the exact dependencies the algorithm should respect or break. The approach relies on a new Langevin-type algorithm that operates on a modified energy function, where parts of the latent variables are averaged over samples from earlier iterations of the Markov chain. This way, statistical dependencies can be broken in a controlled way, allowing the chain to mix faster. This scheme can be further modified in a"dropout"manner, leading to even more scalability. We test our scheme for ResNet-20 on CIFAR-10, SVHN, and FMNIST. In all cases, we find improvements in convergence speed and/or final accuracy compared to SG-MCMC and VI.