Variable-length categoryn-gram language models

Variable-length categoryn-gram language models
复制标题

变长categoryn-gram语言模型

DOI:
10.1006/csla.1998.0115
复制
发表时间:
1999
期刊:
Comput. Speech Lang.
影响因子:
--
通讯作者:
P. Woodland
P. Woodland
中科院分区:
--
文献类型:
--
作者:
T. Niesler;P. Woodland

文献摘要

被引文献

相似文献

本文提出了一种基于n元语法的语言模型。根据预测质量的改进估计,有选择地增加每个字符串的长度。这允许控制模型大小,同时在有利于性能时包括更长范围的依赖关系。类别的选择,以对应的词性分类,在出价开发priorita语法信息。为了解释不同的语法功能,语言模型允许单词属于多个类别,并且隐含地涉及可以用于标记新文本的统计标记操作。基于类别的模型的内在泛化导致稀疏数据集的良好性能。然而,随着训练材料数量的增加,基于单词的n-gram提供了上级平均性能。然而,类别模型继续为训练集中不存在的词元组提供更好的预测。因此,提出了一种允许在退避框架内组合这两种方法的方法。LOB,Switchboard和华尔街日报语料库的实验表明,该技术大大改善了稀疏训练集的语言模型的困惑,并提供了显着改善的大小与性能的权衡时,与标准的三元模型。
Abstract This paper presents a language model based onn-grams of word groups (categories). The length of eachn-gram is increased selectively according to an estimate of the resulting improvement in predictive quality. This allows the model size to be controlled while including longer-range dependencies when these benefit performance. The categories are chosen to correspond to part-of-speech classifications in a bid to exploita priorigrammatical information. To account for different grammatical functions, the language model allows words to belong to multiple categories, and implicitly involves a statistical tagging operation which may be used to label new text. Intrinsic generalization by the category-based model leads to good performance with sparse data sets. However word-basedn-grams deliver superior average performance as the amount of training material increases. Nevertheless, the category model continues to supply better predictions for wordn-tuples not present in the training set. Consequently, a method allowing the two approaches to be combined within a backoff framework is presented. Experiments with the LOB, Switchboard and Wall Street Journal corpora demonstrate that this technique greatly improves language model perplexities for sparse training sets, and offers significantly improved size vs. performance tradeoffs when compared with standard trigram models.