Large scale discriminative training of hidden Markov models for speech recognition

Large scale discriminative training of hidden Markov models for speech recognition
复制标题

DOI:
10.1006/csla.2001.0182
复制
发表时间:
2002-01-01
影响因子:
4.3
通讯作者:
Povey, D
Povey, D
中科院分区:
计算机科学3区
文献类型:
--
作者:
Woodland, PC;Povey, D

文献摘要

被引文献

相似文献

本文描述并大规模评价了基于混合高斯型隐马尔可夫模型(HMM)的大词汇量语音识别系统的格型区分训练框架。本文主要研究最大互信息估计(MMIE)准则,该准则已被用于训练HMM系统,用于会话电话语音转录,使用了长达265小时的训练数据。这些实验代表了作者所知道的语音识别区分训练技术的最大规模应用。给出了与扩展的Baum-Welch算法一起使用的基于MMIE格子的实现的细节,这使得对这样的大型系统的训练在计算上是可行的。讨论了使用声学尺度和弱化语言模型来改进泛化的技术。与使用最大似然估计(MLE)训练的最佳系统相比,整体技术已经允许估计三音素和五音素HMM参数,这使得会话电话语音转录的单词错误率显着降低。这与之前的一些研究形成了鲜明对比,这些研究得出的结论是,对于最困难的大词汇量语音识别任务,使用区别性训练几乎没有什么好处。基于格型MMIE的鉴别训练方案的性能也优于帧鉴别技术。研究了基于网格的MMIE训练方案的各种特性,包括不同的网格处理策略(完全搜索和精确匹配)的比较以及网格大小对性能的影响。此外,还给出了一种基于MMIE和MLE目标函数的线性内插的方案,以减少过度训练的危险。结果表明,用MMIE训练的隐马尔可夫模型与用最大似然线性回归(MLLR)训练的隐马尔可夫模型的模型自适应效果相当。这使得MMIE训练的HMM可以直接集成到复杂的多通道系统中,用于转录对话电话语音,并有助于我们的MMIE训练系统在2000和2001年的NIST Hub5评估中提供最低的单词错误率。(C)2002年学术出版社。
This paper describes, and evaluates on a large scale, the lattice based framework for discriminative training of large vocabulary speech recognition systems based on Gaussian mixture hidden Markov models (HMMs). This paper concentrates on the maximum mutual information estimation (MMIE) criterion which has been used to train HMM systems for conversational telephone speech transcription using up to 265 hours of training data. These experiments represent the largest-scale application of discriminative training techniques for speech recognition of which the authors are aware. Details are given of the MMIE lattice-based implementation used with the extended Baum-Welch algorithm, which makes training of such large systems computationally feasible. Techniques for improving generalization using acoustic scaling and weakened language models are discussed. The overall technique has allowed the estimation of triphone and quinphone HMM parameters which has led to significant reductions in word error rate for the transcription of conversational telephone speech relative to our best systems trained using maximum likelihood estimation (MLE). This is in contrast to some previous studies, which have concluded that there is little benefit in using discriminative training for the most difficult large vocabulary speech recognition tasks. The lattice MMIE-based discriminative training scheme is also shown to out-perform the frame discrimination technique. Various properties of the lattice-based MMIE training scheme are investigated including comparisons of different lattice processing strategies (full search and exact-match) and the effect of lattice size on performance. Furthermore a scheme based on the linear interpolation of the MMIE and MLE objective functions is shown to reduce the danger of over-training. It is shown that HMMs trained with MMIE benefit as much as MLE-trained HMMs from applying model adaptation using maximum likelihood linear regression (MLLR). This has allowed the straightforward integration of MMIE-trained HMMs into complex multi-pass systems for transcription of conversational telephone speech and has contributed to our MMIE-trained systems giving the lowest word error rates in both the 2000 and 2001 NIST Hub5 evaluations. (C) 2002 Academic Press.