Context-Dependent Token-Wise Variational Autoencoder for Topic Modeling

Context-Dependent Token-Wise Variational Autoencoder for Topic Modeling
复制标题

DOI:
10.1007/978-3-030-51253-8_6
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
Tomonari Masada
Tomonari Masada
中科院分区:
其他
文献类型:
--
作者:
Tomonari Masada

文献摘要

相似文献

本文提出了一种新的主题模型变分自动编码器(VAE)。贝叶斯模型的变分推理(VI)通过最大化观测值的对数边际似然的下限来近似真实的后验分布。我们可以通过使用称为编码器的神经网络来获得近似后验的参数来实现VI作为VAE。我们的贡献是三方面的。首先,我们边缘化的每文档的主题概率以下的建议Mimno等人。这种边缘化还没有被认为是在现有的VAE建议的主题建模,以我们所知。其次,在边缘化主题概率之后,我们需要近似token-wise主题分配的后验概率。然而,这个后验是分类的,因此不能用连续分布(如高斯分布)来近似。因此,我们采用Gumbel-softmax技巧。第三,虽然我们可以使用Gumbel-softmax技巧对令牌式主题分配进行采样,但我们应该考虑文档范围的上下文信息以获得更好的近似。因此,我们向我们的编码器网络提供了一个令牌信息和文档范围信息的级联,其中前者被实现为单词嵌入,后者作为单词嵌入的文档范围均值。实验结果表明,我们的VAE改善了现有的VAE建议的一半的数据集的困惑或归一化成对互信息(NPMI)。
This paper proposes a new variational autoencoder (VAE) for topic models. The variational inference (VI) for Bayesian models approximates the true posterior distribution by maximizing a lower bound of the log marginal likelihood of observations. We can implement VI as VAE by using a neural network called encoder to obtain parameters of approximate posterior. Our contribution is three-fold. First, we marginalize out per-document topic probabilities by following the proposal by Mimno et al. This marginalizing out has not been considered in the existing VAE proposals for topic modeling to the best of our knowledge. Second, after marginalizing out topic probabilities, we need to approximate the posterior probabilities of token-wise topic assignments. However, this posterior is categorical and thus cannot be approximated by continuous distributions like Gaussian. Therefore, we adopt the Gumbel-softmax trick. Third, while we can sample token-wise topic assignments with the Gumbel-softmax trick, we should consider document-wide contextual information for a better approximation. Therefore, we feed to our encoder network a concatenation of token-wise information and document-wide information, where the former is implemented as word embedding and the latter as the document-wide mean of word embeddings. The experimental results showed that our VAE improved the existing VAE proposals for a half of the data sets in terms of perplexity or of normalized pairwise mutual information (NPMI).