Marginalized Stochastic Natural Gradients for Black-Box Variational Inference

Marginalized Stochastic Natural Gradients for Black-Box Variational Inference
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Geng Ji;Debora Sujono;Erik B. Sudderth
Geng Ji;Debora Sujono;Erik B. Sudderth
中科院分区:
其他
文献类型:
--
作者:
Geng Ji;Debora Sujono;Erik B. Sudderth

文献摘要

相似文献

黑盒变分推理算法使用随机采样来分析不同的统计模型,如用概率编程语言表达的模型,而没有模型特定的推导。虽然流行的评分函数估计计算无偏梯度估计,它的方差往往是不可接受的大,特别是在离散的潜变量模型。我们提出了一个随机自然梯度估计,是广泛适用的和无偏的,但提高了效率,利用曲率的变分界限,并证明减少方差边缘化离散潜变量。我们的边缘化的随机自然梯度有有趣的连接到经典的坐标上升变分推理,但允许并行更新的变分参数,并提供上级收敛保证相对于天真的蒙特卡罗近似。我们将我们的方法与概率编程语言Pyro相结合,并评估文档,图像,网络和众包的真实模型。与分数函数估计相比,我们需要更少的蒙特卡罗样本和一致的收敛速度更快的数量级。
Black-box variational inference algorithms use stochastic sampling to analyze diverse statistical models, like those expressed in probabilistic programming languages, without model-specific derivations. While the popular score-function estimator computes unbiased gradient estimates, its variance is often unacceptably large, especially in models with discrete latent variables. We propose a stochastic natural gradient estimator that is as broadly applicable and unbiased, but improves efficiency by exploiting the curvature of the variational bound, and provably reduces variance by marginalizing discrete latent variables. Our marginalized stochastic natural gradients have intriguing connections to classic coordinate ascent variational inference, but allow parallel updates of variational parameters, and provide superior convergence guarantees relative to naive Monte Carlo approximations. We integrate our method with the probabilistic programming language Pyro and evaluate real-world models of documents, images, networks, and crowd-sourcing. Compared to score-function estimators, we require far fewer Monte Carlo samples and consistently convergence orders of magnitude faster.