Langevin Autoencoders for Learning Deep Latent Variable Models

Langevin Autoencoders for Learning Deep Latent Variable Models
复制标题

DOI:
10.48550/arxiv.2209.07036
复制
发表时间:
2022-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Shohei Taniguchi;Yusuke Iwasawa;Wataru Kumagai;Yutaka Matsuo
Shohei Taniguchi;Yusuke Iwasawa;Wataru Kumagai;Yutaka Matsuo
中科院分区:
其他
文献类型:
--
作者:
Shohei Taniguchi;Yusuke Iwasawa;Wataru Kumagai;Yutaka Matsuo

文献摘要

相似文献

马尔可夫链蒙特卡罗(MCMC),如朗之万动力学,是有效的近似棘手的分布。然而,它的使用是有限的背景下,深潜变量模型由于昂贵的数据点采样迭代和缓慢的收敛。本文提出了摊销朗之万动力学(ALD),其中数据逐点MCMC迭代完全替换为更新的编码器,将观测映射到潜在变量。这种摊销实现了有效的后验采样,而无需逐数据点迭代。尽管它的效率,我们证明了ALD是有效的MCMC算法,其马尔可夫链的目标后验作为一个平稳分布在温和的假设。基于ALD,我们还提出了一个新的深层潜变量模型命名为Langevin自动编码器(LAE)。有趣的是,LAE可以通过稍微修改传统的自动编码器来实现。使用多个合成数据集,我们首先验证ALD可以正确地从目标后验子中获取样本。我们还评估了LAE的图像生成任务,并表明我们的LAE可以优于现有的方法的基础上变分推理,如变分自动编码器,和其他MCMC为基础的方法在测试的可能性。
Markov chain Monte Carlo (MCMC), such as Langevin dynamics, is valid for approximating intractable distributions. However, its usage is limited in the context of deep latent variable models owing to costly datapoint-wise sampling iterations and slow convergence. This paper proposes the amortized Langevin dynamics (ALD), wherein datapoint-wise MCMC iterations are entirely replaced with updates of an encoder that maps observations into latent variables. This amortization enables efficient posterior sampling without datapoint-wise iterations. Despite its efficiency, we prove that ALD is valid as an MCMC algorithm, whose Markov chain has the target posterior as a stationary distribution under mild assumptions. Based on the ALD, we also present a new deep latent variable model named the Langevin autoencoder (LAE). Interestingly, the LAE can be implemented by slightly modifying the traditional autoencoder. Using multiple synthetic datasets, we first validate that ALD can properly obtain samples from target posteriors. We also evaluate the LAE on the image generation task, and show that our LAE can outperform existing methods based on variational inference, such as the variational autoencoder, and other MCMC-based methods in terms of the test likelihood.