VAE-KRnet and its applications to variational Bayes

VAE-KRnet and its applications to variational Bayes
复制标题

DOI:
10.4208/cicp.oa-2021-0087
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
X. Wan;Shuangqing Wei
X. Wan;Shuangqing Wei
中科院分区:
其他
文献类型:
--
作者:
X. Wan;Shuangqing Wei

文献摘要

相似文献

在这项工作中,我们提出了一个生成模型的密度估计,称为VAE-KRnet,它结合了规范变分自动编码器(VAE)与我们最近开发的基于流的生成模型,称为KRnet。VAE被用作一种降维技术来捕获潜在空间,KRnet被用于对潜在变量的分布进行建模。使用数据和潜变量之间的线性模型,我们表明VAE-KRnet可以比规范VAE更有效和鲁棒。作为应用,我们将VAE-KRnet应用于变分贝叶斯来近似后验。变分贝叶斯方法通常基于模型和后验之间的Kullback-Leibler(KL)散度的最小化,如果模型能力不够强,这通常会低估方差。然而,对于高维分布,构建准确的模型是非常具有挑战性的,因为为了效率,通常需要额外的假设,例如,平均场方法假设维度之间相互独立。当维数相对较小时,KRnet可以用来有效地近似原始随机变量的后验。对于高维情况,我们考虑VAE-KRnet与降维。为了减轻低估的方差,我们包括最大化的互信息之间的潜在的随机变量和原始的,当寻求一个近似的分布相对于KL分歧。数值实验证明了该模型的有效性。
In this work, we have proposed a generative model for density estimation, called VAE-KRnet, which combines the canonical variational autoencoder (VAE) with our recently developed flow-based generative model, called KRnet. VAE is used as a dimension reduction technique to capture the latent space, and KRnet is used to model the distribution of the latent variables. Using a linear model between the data and the latent variables, we show that VAE-KRnet can be more effective and robust than the canonical VAE. As an application, we apply VAE-KRnet to variational Bayes to approximate the posterior. The variational Bayes approaches are usually based on the minimization of the Kullback-Leibler (KL) divergence between the model and the posterior, which often underestimates the variance if the model capability is not sufficiently strong. However, for high-dimensional distributions, it is very challenging to construct an accurate model since extra assumptions are often needed for efficiency, e.g., the mean-field approach assumes mutual independence between dimensions. When the number of dimensions is relatively small, KRnet can be used to approximate the posterior effectively with respect to the original random variable. For high-dimensional cases, we consider VAE-KRnet to incorporate with the dimension reduction. To alleviate the underestimation of the variance, we include the maximization of the mutual information between the latent random variable and the original one when seeking an approximate distribution with respect to the KL divergence. Numerical experiments have been presented to demonstrate the effectiveness of our model.