Structured and Efficient Variational Deep Learning with Matrix Gaussian Posteriors

Structured and Efficient Variational Deep Learning with Matrix Gaussian Posteriors
复制标题

DOI:
--
复制
发表时间:
2016-03
期刊:
--
影响因子:
--
通讯作者:
Christos Louizos;M. Welling
Christos Louizos;M. Welling
中科院分区:
其他
文献类型:
--
作者:
Christos Louizos;M. Welling

文献摘要

被引文献

相似文献

我们介绍了一个变分贝叶斯神经网络的参数是通过随机矩阵的概率分布。具体来说,我们采用了一个矩阵变量高斯(Gupta & Nagar,1999)参数后验分布,其中我们明确地对每个层的输入和输出维度之间的协方差进行建模。此外,通过近似协方差矩阵,我们可以实现一种更有效的方式来表示这些相关性,这也比完全因子分解的参数后验更便宜。我们进一步表明,使用“局部重新度量技巧”(Kingma等人,2015)在这个后验分布上,我们得到了每层隐藏单元的高斯过程(Rasmussen,2006)解释,并且我们与(Gal & Ghahramani,2015)类似,提供了与深度高斯过程的连接。我们继续利用这种二元性,并在我们的模型中加入“伪数据”(Snelson & Ghahramani,2005),这反过来又允许更有效的后验采样,同时保持原始模型的属性。通过大量的实验验证了该方法的有效性。
We introduce a variational Bayesian neural network where the parameters are governed via a probability distribution on random matrices. Specifically, we employ a matrix variate Gaussian (Gupta & Nagar, 1999) parameter posterior distribution where we explicitly model the covariance among the input and output dimensions of each layer. Furthermore, with approximate covariance matrices we can achieve a more efficient way to represent those correlations that is also cheaper than fully factorized parameter posteriors. We further show that with the "local reprarametrization trick" (Kingma et al., 2015) on this posterior distribution we arrive at a Gaussian Process (Rasmussen, 2006) interpretation of the hidden units in each layer and we, similarly with (Gal & Ghahramani, 2015), provide connections with deep Gaussian processes. We continue in taking advantage of this duality and incorporate "pseudo-data" (Snelson & Ghahramani, 2005) in our model, which in turn allows for more efficient posterior sampling while maintaining the properties of the original model. The validity of the proposed approach is verified through extensive experiments.