Generalized Algorithms for Constructing Statistical Language Models

Generalized Algorithms for Constructing Statistical Language Models
复制标题

构建统计语言模型的通用算法

DOI:
10.3115/1075096.1075102
复制
发表时间:
2003
期刊:
Annual Meeting of the Association for Computational Linguistics
影响因子:
--
通讯作者:
Brian Roark
Brian Roark
中科院分区:
--
文献类型:
--
作者:
Cyril Allauzen;Mehryar Mohri;Brian Roark

文献摘要

被引文献

相似文献

最近的文本和语音处理应用,如语音挖掘,提出了与语言模型构建相关的新的和更普遍的问题。我们提出并详细描述了几个新的和有效的算法来解决这些更普遍的问题,并报告了实验结果,证明了它们的实用性。我们给出了一种算法,用于有效地计算语音识别器或任意加权自动机输出的字格中任意序列的期望计数;描述一种通过加权自动机创建n-gram语言模型精确表示的新技术,该技术的大小适用于离线使用,即使词汇量约为500,000个单词,n-gram阶数n = 6;并提出一种简单而通用的技术,用于构建基于类的语言模型,该模型允许每个类表示任意加权自动机。我们的算法和技术的一个有效实现已经被整合到一个用于语言建模的通用软件库中,GRM库,其中包括许多其他文本和语法处理功能。
Recent text and speech processing applications such as speech mining raise new and more general problems related to the construction of language models. We present and describe in detail several new and efficient algorithms to address these more general problems and report experimental results demonstrating their usefulness. We give an algorithm for computing efficiently the expected counts of any sequence in a word lattice output by a speech recognizer or any arbitrary weighted automaton; describe a new technique for creating exact representations of n-gram language models by weighted automata whose size is practical for offline use even for a vocabulary size of about 500,000 words and an n-gram order n = 6; and present a simple and more general technique for constructing class-based language models that allows each class to represent an arbitrary weighted automaton. An efficient implementation of our algorithms and techniques has been incorporated in a general software library for language modeling, the GRM Library, that includes many other text and grammar processing functionalities.