Heuristic Pretraining for Topic Models

Heuristic Pretraining for Topic Models
复制标题

DOI:
10.1007/978-3-319-19066-2_34
复制
发表时间:
2015-06
期刊:
--
影响因子:
--
通讯作者:
Tomonari Masada;A. Takasu
Tomonari Masada;A. Takasu
中科院分区:
其他
文献类型:
--
作者:
Tomonari Masada;A. Takasu

文献摘要

相似文献

本文提出了一种启发式的主题模型预训练方法。虽然我们在这里考虑了潜在狄利克雷分配(LDA),但我们的预训练可以应用于其他主题模型。基本上,我们使用折叠吉布斯采样(CGS)来更新潜变量。然而,在CGS的每次迭代之后,我们将潜变量视为可观察的,并在它们之上构造另一个LDA,我们称之为LDA over LDA(LoL)。然后,我们执行以下两种类型的更新:通过CGS更新LoL中的潜变量,以及基于LoL中潜变量的先前更新结果更新LDA中的潜变量。我们执行一个迭代的CGS LDA和上述两种类型的更新交替只为一个小的,早期的一部分推理。也就是说,所提出的方法被用作apretraining。预训练阶段之后是用于LDA的CGS的通常迭代。评价实验表明,我们的预训练可以改善测试集的困惑。
This paper provides a heuristic pretraining for topic models. While we consider latent Dirichlet allocation (LDA) here, our pretraining can be applied to other topic models. Basically, we use collapsed Gibbs sampling (CGS) to update the latent variables. However, after every iteration of CGS, we regard the latent variables as observable and construct another LDA over them, which we callLDA over LDA (LoL). We then perform the following two types of updates: the update of the latent variables in LoL by CGS and the update of the latent variables in LDA based on the result of the preceding update of the latent variables in LoL. We perform one iteration of CGS for LDA and the above two types of updates alternately only for a small, earlier part of the inference. That is, the proposed method is used as apretraining. The pretraining stage is followed by the usual iterations of CGS for LDA. The evaluation experiment shows that our pretraining can improve test set perplexity.