Latent Dirichlet Allocation

Latent Dirichlet Allocation
复制标题

DOI:
10.7766/orbit.v1.2.44
复制
发表时间:
2001-01
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
D. Blei;Andrew Y. Ng;Michael I. Jordan
D. Blei;Andrew Y. Ng;Michael I. Jordan
中科院分区:
其他
文献类型:
--
作者:
D. Blei;Andrew Y. Ng;Michael I. Jordan

文献摘要

被引文献

相似文献

我们提出了一种用于文本和其他离散数据集合的生成模型,该模型概括或改进了以前的几种模型,包括朴素贝叶斯/unigram,unigram的混合[6]和霍夫曼的方面模型,也称为概率潜在语义索引(pLSI)[3]。在文本建模的背景下,我们的模型假定每个文档都是作为主题的混合物生成的,其中连续值混合比例作为潜在的Dirichlet随机变量分布。通过变分算法有效地进行推理和学习。我们目前的实证结果,这个模型的应用程序的问题,在文本建模,协同过滤和文本分类。
We propose a generative model for text and other collections of discrete data that generalizes or improves on several previous models including naive Bayes/unigram, mixture of unigrams [6], and Hofmann's aspect model , also known as probabilistic latent semantic indexing (pLSI) [3]. In the context of text modeling, our model posits that each document is generated as a mixture of topics, where the continuous-valued mixture proportions are distributed as a latent Dirichlet random variable. Inference and learning are carried out efficiently via variational algorithms. We present empirical results on applications of this model to problems in text modeling, collaborative filtering, and text classification.