The Development of Topic Models in Natural Language Processing

The Development of Topic Models in Natural Language Processing
复制标题

DOI:
10.3724/sp.j.1016.2011.01423
复制
发表时间:
2011
期刊:
Chinese Journal of Computers
影响因子:
--
通讯作者:
Wang Hou
Wang Hou
中科院分区:
其他
文献类型:
--
作者:
Wang Hou

文献摘要

被引文献

相似文献

主题模型是自然语言处理领域的一个研究热点,在该领域中,主题被看作是术语的概率分布,主题模型利用术语在文档级的共现来提取语义主题,并将位于术语空间的文档转换为主题空间的文档,从而获得文档的低维表示。本文从主题模型的起源--潜在语义索引(LSI)出发,描述了主题模型发展的基础工作pLSI和LDA,重点讨论了它们之间的关系。LDA作为一种生成式模型,可以很容易地扩展到其他模型中。本文对由LDA派生的主题模型进行了简单的分类,并介绍了每类主题模型的代表模型。此外,本文还讨论了潜在语义索引的概念,并对潜在语义索引中的主题模型进行了分类。分析了主题模型参数估计中的EM算法,有助于理解主题模型开发过程中作品之间的关系。
Topic models are receiving extensive attention in natural language processing.In this field,a topic is regarded as probabilistic distribution of terms.Topic models extract semantic topics using co-occurrence of terms in document level,and are used to transform documents locating in term space to the ones in topic space,obtaining the low dimensional representation of documents. This paper starts from Latent Semantic Indexing(LSI),the origin of topic models,and describes pLSI and LDA,the fundamental works in the development of topic models,with focus on the relationship among these works.As a generative model,LDA can be easily extended to other models.This paper makes a simple categorization on topic models derived from LDA,and representative models of each category are introduced.Furthermore,EM algorithms in parameter estimation of topic models are analyzed,which help to understand the relationship of works during the development of topic models.