Topic-based language models using Dirichlet Mixtures
Topic-based language models using Dirichlet Mixtures
复制标题
DOI:
10.1002/scj.20629
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
Kugatsu Sadamitsu;T. Mishina;Mikio Yamamoto
中科院分区:
文献类型:
--
作者:
Kugatsu Sadamitsu;T. Mishina;Mikio Yamamoto
We propose a generative text model using Dirichlet Mixtures as a distribution for parameters of a multinomial distribution, whose compound distribution is Polya Mixtures, and show that the model exhibits high performance in application to statistical language models. In this paper, we discuss some methods for estimating parameters of Dirichlet Mixtures and for estimating the expectation values of the a posteriori distribution needed for adaptation, and then compare them with two previous text models. The first conventional model is the Mixture of Unigrams, which is often used for incorporating topics into statistical language models. The second one is LDA (Latent Dirichlet Allocation), a typical generative text model. In an experiment using document probabilities and dynamic adaptation of n-gram models for newspaper articles, we show that the proposed model, in comparison with the two previous models, can achieve a lower perplexity at low mixture numbers. © 2007 Wiley Periodicals, Inc. Syst Comp Jpn, 38(12): 76–85, 2007; Published online in Wiley InterScience (www.interscience.wiley.com). DOI 10.1002/scj.20629