Finding scientific topics

Finding scientific topics
复制标题

DOI:
10.1073/pnas.0307752101
复制
发表时间:
2004-04-06
影响因子:
11.1
通讯作者:
Steyvers, M
Steyvers, M
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Griffiths, TL;Steyvers, M

文献摘要

被引文献

相似文献

识别文档内容的第一步是确定文档涉及哪些主题。我们描述了一个生成模型的文件,介绍了Blei,Ng和约旦[Blei,D。M.,Ng,A. Y. & Jordan,M. I.(2003)机器学习。Res.3,993-1022],其中通过在主题上选择分布并且然后从根据该分布选择的主题中选择文档中的每个词来生成每个文档。然后,我们提出了一个马尔可夫链蒙特卡罗算法在这个模型中的推理。我们使用这个算法来分析摘要从PNAS使用贝叶斯模型选择建立的主题数。我们发现,提取的主题捕捉有意义的结构的数据,与类指定的作者提供的文章,并概述了进一步的应用,这种分析,包括确定“热门话题”通过检查时间动态和标记摘要,以说明语义内容。
A first step in identifying the content of a document is determining which topics that document addresses. We describe a generative model for documents, introduced by Blei, Ng, and Jordan [Blei, D. M., Ng, A. Y. & Jordan, M. I. (2003) J. Machine Learn. Res. 3, 993-1022], in which each document is generated by choosing a distribution over topics and then choosing each word in the document from a topic selected according to this distribution. We then present a Markov chain Monte Carlo algorithm for inference in this model. We use this algorithm to analyze abstracts from PNAS by using Bayesian model selection to establish the number of topics. We show that the extracted topics capture meaningful structure in the data, consistent with the class designations provided by the authors of the articles, and outline further applications of this analysis, including identifying "hot topics" by examining temporal dynamics and tagging abstracts to illustrate semantic content.