Distributed Bayesian Learning with Stochastic Natural Gradient Expectation Propagation and the Posterior Server

Distributed Bayesian Learning with Stochastic Natural Gradient Expectation Propagation and the Posterior Server
复制标题

DOI:
--
复制
发表时间:
2015-12
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Leonard Hasenclever;Stefan Webb;Thibaut Lienart;S. Vollmer;Balaji Lakshminarayanan;C. Blundell;Y. Teh
Leonard Hasenclever;Stefan Webb;Thibaut Lienart;S. Vollmer;Balaji Lakshminarayanan;C. Blundell;Y. Teh
中科院分区:
其他
文献类型:
--
作者:
Leonard Hasenclever;Stefan Webb;Thibaut Lienart;S. Vollmer;Balaji Lakshminarayanan;C. Blundell;Y. Teh

文献摘要

被引文献

相似文献

本文对贝叶斯机器学习算法做了两个贡献。首先,我们提出了随机自然梯度期望传播(SNEP),一种新的替代期望传播(EP),一个流行的变分推理算法。SNEP是一种黑盒变分算法,因为它不需要对感兴趣的分布进行任何简化假设,除了用于估计EP倾斜分布的矩的一些Monte Carlo采样器的存在之外。此外,相对于EP没有收敛保证,SNEP可以被证明是收敛的,即使使用蒙特卡罗矩估计。其次,我们提出了一种新的分布式贝叶斯学习架构,我们称之为后验服务器。后验服务器允许在数据集以分布式方式存储在集群中的情况下进行可扩展和鲁棒的贝叶斯学习,每个计算节点包含不相交的数据子集。在每个计算节点上运行一个独立的Monte Carlo采样器,仅直接访问本地数据子集,但其目标是在整个集群中给定所有数据的情况下近似全局后验分布。这是通过使用SNEP的分布式异步实现跨集群传递消息来实现的。我们展示了SNEP和后验服务器对分布式贝叶斯学习的逻辑回归和神经网络。保留字:分布式学习,大规模学习,深度学习,贝叶斯学习,变分推理,期望传播,随机近似,自然梯度,马尔可夫链蒙特卡罗,参数服务器,后验服务器。
This paper makes two contributions to Bayesian machine learning algorithms. Firstly, we propose stochastic natural gradient expectation propagation (SNEP), a novel alternative to expectation propagation (EP), a popular variational inference algorithm. SNEP is a black box variational algorithm, in that it does not require any simplifying assumptions on the distribution of interest, beyond the existence of some Monte Carlo sampler for estimating the moments of the EP tilted distributions. Further, as opposed to EP which has no guarantee of convergence, SNEP can be shown to be convergent, even when using Monte Carlo moment estimates. Secondly, we propose a novel architecture for distributed Bayesian learning which we call the posterior server. The posterior server allows scalable and robust Bayesian learning in cases where a data set is stored in a distributed manner across a cluster, with each compute node containing a disjoint subset of data. An independent Monte Carlo sampler is run on each compute node, with direct access only to the local data subset, but which targets an approximation to the global posterior distribution given all data across the whole cluster. This is achieved by using a distributed asynchronous implementation of SNEP to pass messages across the cluster. We demonstrate SNEP and the posterior server on distributed Bayesian learning of logistic regression and neural networks. Keywords: Distributed Learning, Large Scale Learning, Deep Learning, Bayesian Learn- ing, Variational Inference, Expectation Propagation, Stochastic Approximation, Natural Gradient, Markov chain Monte Carlo, Parameter Server, Posterior Server.