Improvements of Katz K Mixture Model

Improvements of Katz K Mixture Model
复制标题

Katz K混合模型的改进

DOI:
10.5715/jnlp.12.5_131
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
Kyoji Umemura
Kyoji Umemura
中科院分区:
--
文献类型:
--
作者:
Yinghui Xu;Kyoji Umemura

文献摘要

被引文献

相似文献

拟合经验词分布以及负二项式的一个更简单的分布是 Katz K 混合模型。在 K 混合模型中,基本假设是给定词的重复条件概率由与已发生次数无关的常数衰减因子决定。然而,重复出现的概率通常低于出现次数很少的包含内容的词的常数衰减因子。为了解决 K 混合模型的这一缺陷,深入研究了 K 混合模型的不足之处。探索了重复的条件概率、衰减因素及其对建模术语分布的影响。根据本研究的结果,似乎可以使用分布的两端来拟合模型。即,不仅可以在单词实例较少时使用文档频率,还可以使用尾部概率(文档频率的累积)。一个单词的少数实例的文档频率和大型实例的尾部概率通常相对容易凭经验估计。因此,我们提出了一种改进 K 混合模型的有效方法,其中衰减因子是根据文档中单词的实例数量由函数插值的两个可能的衰减因子的组合。结果表明,所提出的模型可以生成统计上显着的更好的频率估计,特别是文档中具有两个实例的单词的频率估计。此外,表明该方法的优点将成为在两种情况下更为明显,对经常使用的内容承载词的术语分布进行建模,以及对具有广泛文档长度的语料库进行术语分布建模。
A simpler distribution that fits empirical word distribution about as well as a negative binomial is the Katz K mixture.In the K mixture model, the basic assumption is that the conditional probabilities of repeats for a given word are determined by a constant decay factor that is independent of the number of occurrences which have taken place.However, the probabilities of the repeat occurrences are generally lower than the constant decay factor for the content-bearing words with few occurrences that have taken place.To solve this deficiency of the K mixture model, in-depth exploration of the characteristics of the conditional probabilities of repetitions, decay factors and their influences on modeling term distributions was conducted.Based on the results of this study, it appears that both ends of the distribution can be used to fit models.That is, not only can document frequencies be used when the instances of a word are few, but also tail probabilities (the accumulation of document frequencies). Both document frequencies for few instances of a word and tail probabilities for large instances are often relatively easy to estimate empirically.Therefore, we propose an effective approach for improving the K mixture model, where the decay factor is the combination of two possible decay factors interpolated by a function depending on the number of instances of a word in a document.Results show that the proposed model can generate a statistically significant better estimation of frequencies, especially the frequency estimation for a word with two instances in a document.In addition, it is shown that the advantages of this approach will become more evident in two cases, modeling the term distribution for the frequently used content-bearing word and modeling the term distribution for a corpus with a wide range of document length.