On structuring probabilistic dependences in stochastic language modelling

On structuring probabilistic dependences in stochastic language modelling
复制标题

DOI:
10.1006/csla.1994.1001
复制
发表时间:
1994
期刊:
Comput. Speech Lang.
影响因子:
--
通讯作者:
H. Ney;U. Essen;Reinhard Kneser
H. Ney;U. Essen;Reinhard Kneser
中科院分区:
其他
文献类型:
--
作者:
H. Ney;U. Essen;Reinhard Kneser

文献摘要

被引文献

相似文献

摘要在本文中,我们从将合适的结构引入条件概率分布的角度研究了随机语言建模的问题。这些分布的任务是通过查看M甚至所有前任单词来预测新词的概率。常规的方法是将M限制为1或2,并以线性方式插入带有Unigram模型的最终的Bigram和Trigram模型。但是,还有许多其他结构可用于建模前任单词和要预测的单词之间的概率依赖性。本文考虑的结构是:非线性插值作为线性插值的替代方案;单词历史和单词的等效类;缓存内存和单词关联。为了对非线性和线性插值参数进行最佳估计,系统地使用了剩下的一个方法。为了确定BigRAM模型中的单词等价类,已经对自动聚类过程进行了调整。为了捕获长距离依赖,我们考虑了各种逐字依赖的模型。缓存模型可能被视为一种特殊的自我关联类型。提供了两个文本数据库,一个德国数据库和一个英语数据库的实验结果。
Abstract In this paper, we study the problem of stochastic language modelling from the viewpoint of introducing suitable structures into the conditional probability distributions. The task of these distributions is to predict the probability of a new word by looking at M or even all predecessor words. The conventional approach is to limit M to 1 or 2 and to interpolate the resulting bigram and trigram models with a unigram model in a linear fashion. However, there are many other structures that can be used to model the probabilistic dependences between the predecessor word and the word to be predicted. The structures considered in this paper are: nonlinear interpolation as an alternative to linear interpolation; equivalence classes for word histories and single words; cache memory and word associations. For the optimal estimation of nonlinear and linear interpolation parameters, the leaving-one-out method is systematically used. For the determination of word equivalence classes in a bigram model, an automatic clustering procedure has been adapted. To capture long-distance dependences, we consider various models for word-by-word dependences; the cache model may be viewed as a special type of self-association. Experimental results are presented for two text databases, a Germany database and an English database.