AmLDA: A Non-VAE Neural Topic Model

AmLDA: A Non-VAE Neural Topic Model
复制标题

AmLDA:非 VAE 神经主题模型

DOI:
10.1007/978-3-031-04447-2_19
复制
发表时间:
2022
期刊:
Springer Communications in Computer and Information Science
影响因子:
--
通讯作者:
Tomonari MASADA
Tomonari MASADA
中科院分区:
--
文献类型:
--
作者:
Sasaki Toru;Masada Tomonari;Tomonari MASADA

文献摘要

相似文献

贝叶斯模型的变分推断(VI)通过最大化一个称为ELBO的变分下界来逼近真实的后验。本文考虑了潜在狄利克雷分配(LDA)的VI,这是一个研究得很好的贝叶斯模型。LDA的VI最初是作为一种变分期望最大化(VEM)提出的,通过将ELBO的导数设为零来得到更新方程。然而,它的M步需要对整个训练集进行分析,而它的E步需要运行数十次更新。后来提出的随机VI(SVI)改进了VEM,用小批量梯度上升代替了M步。此外,一种用于LDA的变分自动编码器,称为PROLDA,又通过使用摊销编码器,用小批量梯度上升取代了E阶跃。现在我们可以像训练深度神经网络一样训练LDA。然而,PROLDA排除了离散的潜在变量,从而最大化了ELBO公式,而不同于用于LDA的VEM和SVI。因此,PROLDA会遇到零部件崩溃的问题。因此,我们提出了一种新的用于LDA的VI,称为AMLDA。由于AMLDA最大化的ELBO与VEM和SVI最大化的ELBO相同,因此它不会受到组件崩溃的影响。只有参数化不同,因为AMLDA使用分期网络对未知变量进行参数化。评估是在五个大型数据集上进行的。实验结果表明,AMLDA算法与SVI算法具有同样的效果。
The variational inference (VI) for Bayesian models approximates the true posterior by maximizing a variational lower bound called ELBO. This paper considers the VI for the latent Dirichlet allocation (LDA), a well-studied Bayesian model. The VI for LDA was originally proposed as a variational expectation-maximization (VEM), where we obtain the update equations by setting the derivatives of the ELBO equal to zero. However, its M step requires the analysis of the whole training set, and its E step needs to run the update dozens of times. The stochastic VI (SVI) proposed later has improved the VEM by replacing the M step with a minibatch gradient ascent. Further, a variational autoencoder for LDA called ProdLDA in turn has replaced the E step with a minibatch gradient ascent by using an amortized encoder. Now we can train LDA like a deep neural network. However, ProdLDA marginalizes out the discrete latent variables and thus maximizes an ELBO formulated differently from both the VEM and the SVI for LDA. As a result, ProdLDA suffers from the problem of component collapse. Therefore, we propose a new VI for LDA called AmLDA. As AmLDA maximizes the same ELBO as that which both the VEM and the SVI maximize, it does not suffer from component collapse. Only the parameterization differs because AmLDA uses an amortized network for parameterizing unknown variables. The evaluation was performed over five large datasets. The experimental results show that AmLDA is as effective as the SVI.