An empirical study of smoothing techniques for language modeling

An empirical study of smoothing techniques for language modeling
复制标题

DOI:
10.1006/csla.1999.0128
复制
发表时间:
1999-10-01
影响因子:
4.3
通讯作者:
Goodman, J
Goodman, J
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chen, SF;Goodman, J

文献摘要

被引文献

相似文献

我们调查了最广泛使用的算法,平滑模型的语言n-gram建模。然后,我们对几种平滑技术进行了广泛的经验比较,包括Jelinek和美世(1980)、Katz(1987)、Bell、Cleary和维滕(1990)、Ney、埃森和Kneser(1994)以及Kneser和Ney(1995)所描述的平滑技术。我们研究了训练数据大小、训练语料库(例如Brown与Wall Street Journal)、计数截止值和n-gram顺序(bigram与trigram)等因素如何影响这些方法的相对性能,这些性能通过测试数据的交叉熵来衡量。我们发现,这些因素可以显着影响模型的相对性能,其中最重要的因素是训练数据大小。由于没有以前的比较系统地检查这些因素,这是第一次彻底的各种算法的相对性能的表征。此外,我们介绍的方法,详细分析平滑算法的功效,并使用这些技术,我们激励一个新的变化Kneser-Ney平滑,始终优于所有其他算法的评估。最后,结果表明,改进的语言模型平滑导致改善语音识别性能。(C)北京:科学出版社.
We survey the most widely-used algorithms for smoothing models for language n-gram modeling. We then present an extensive empirical comparison of several of these smoothing techniques, including those described by Jelinek and Mercer (1980); Katz (1987); Bell, Cleary and Witten (1990); Ney, Essen and Kneser (1994), and Kneser and Ney (1995). We investigate how factors such as training data size, training corpus (e.g. Brown vs. Wall Street Journal), count cutoffs, and n-gram order (bigram vs. trigram) affect the relative performance of these methods, which is measured through the cross-entropy of test data. We find that these factors can significantly affect the relative performance of models, with the most significant factor being training data size. Since no previous comparisons have examined these factors systematically, this is the first thorough characterization of the relative performance of various algorithms. In addition, we introduce methodologies for analyzing smoothing algorithm efficacy in detail, and using these techniques we motivate a novel variation of Kneser-Ney smoothing that consistently outperforms all other algorithms evaluated. Finally, results showing that improved language model smoothing leads to improved speech recognition performance are presented. (C) 1999 Academic Press.