Topic-based language models using Dirichlet Mixtures

Topic-based language models using Dirichlet Mixtures
复制标题

DOI:
10.1002/scj.20629
复制
发表时间:
2007
期刊:
Systems and Computers in Japan
影响因子:
--
通讯作者:
Kugatsu Sadamitsu;T. Mishina;Mikio Yamamoto
Kugatsu Sadamitsu;T. Mishina;Mikio Yamamoto
中科院分区:
其他
文献类型:
--
作者:
Kugatsu Sadamitsu;T. Mishina;Mikio Yamamoto

文献摘要

相似文献

我们提出了一个生成的文本模型,使用Dirichlet混合作为一个多项分布,其复合分布是Polya混合的参数分布,并表明该模型具有高性能的应用统计语言模型。本文讨论了Dirichlet混合模型的参数估计和适应所需的后验分布的期望值估计的一些方法,并与已有的两个文本模型进行了比较。第一个传统的模型是Unigrams的混合,它通常用于将主题纳入统计语言模型。第二种是LDA(Latent Dirichlet Allocation),一种典型的生成式文本模型。在一个实验中,使用文档概率和动态适应的n-gram模型的报纸文章,我们表明,该模型,与以前的两个模型相比,可以实现较低的困惑在低的混合数。© 2007威利期刊公司Syst Comp Jpn,38(12):76-85,2007;在线发表于Wiley InterScience(www.interscience.wiley.com)。DOI 10.1002/scj.20629
We propose a generative text model using Dirichlet Mixtures as a distribution for parameters of a multinomial distribution, whose compound distribution is Polya Mixtures, and show that the model exhibits high performance in application to statistical language models. In this paper, we discuss some methods for estimating parameters of Dirichlet Mixtures and for estimating the expectation values of the a posteriori distribution needed for adaptation, and then compare them with two previous text models. The first conventional model is the Mixture of Unigrams, which is often used for incorporating topics into statistical language models. The second one is LDA (Latent Dirichlet Allocation), a typical generative text model. In an experiment using document probabilities and dynamic adaptation of n-gram models for newspaper articles, we show that the proposed model, in comparison with the two previous models, can achieve a lower perplexity at low mixture numbers. © 2007 Wiley Periodicals, Inc. Syst Comp Jpn, 38(12): 76–85, 2007; Published online in Wiley InterScience (www.interscience.wiley.com). DOI 10.1002/scj.20629