Scalable Hyperparameter Selection for Latent Dirichlet Allocation

Scalable Hyperparameter Selection for Latent Dirichlet Allocation
复制标题

DOI:
10.1080/10618600.2020.1741378
复制
发表时间:
2020-05
影响因子:
2.4
通讯作者:
Wei Xia;Hani Doss
Wei Xia;Hani Doss
中科院分区:
数学2区
文献类型:
--
作者:
Wei Xia;Hani Doss

文献摘要

被引文献

相似文献

潜在狄利克雷分配(Latent Dirichlet allocation, LDA)是一种被广泛使用的贝叶斯分层模型,用于机器学习中对高维稀疏计数数据(如文本文档)建模。作为一个贝叶斯模型,它包含了一组潜在变量的先验。先验被一些超参数所索引,这些超参数对模型的推理有很大的影响。超参数的理想估计是经验贝叶斯估计,根据定义,它是所有潜在变量积分后数据的边际似然的最大化器。这个估计不能通过分析得到。在实践中,超参数要么以一种特别的方式选择,要么通过一些理论基础薄弱的EM算法的变体来选择。我们提出了一种基于mcmc的全贝叶斯方法来获得超参数的经验贝叶斯估计。并在合成数据和实际数据上与现有方法进行了比较。对比实验表明,该方法所确定的具有超参数的LDA模型优于其他方法估计的具有超参数的LDA模型。本文的补充材料可在网上获得。
Abstract Latent Dirichlet allocation (LDA) is a heavily used Bayesian hierarchical model used in machine learning for modeling high-dimensional sparse count data, for example, text documents. As a Bayesian model, it incorporates a prior on a set of latent variables. The prior is indexed by some hyperparameters, which have a big impact on inference regarding the model. The ideal estimate of the hyperparameters is the empirical Bayes estimate which is, by definition, the maximizer of the marginal likelihood of the data with all the latent variables integrated out. This estimate cannot be obtained analytically. In practice, the hyperparameters are chosen either in an ad-hoc manner, or through some variants of the EM algorithm for which the theoretical basis is weak. We propose an MCMC-based fully Bayesian method for obtaining the empirical Bayes estimate of the hyperparameter. We compare our method with other existing approaches both on synthetic and real data. The comparative experiments demonstrate that the LDA model with hyperparameters specified by our method outperforms models with the hyperparameters estimated by other methods. Supplementary materials for this article are available online.