An N-best candidates-based discriminative training for speech recognition applications

An N-best candidates-based discriminative training for speech recognition applications
复制标题

用于语音识别应用的基于 N 最佳候选的判别训练

DOI:
10.1109/89.260363
复制
发表时间:
1994
期刊:
IEEE Transactions on Speech and Audio Processing
影响因子:
--
通讯作者:
F. Soong
F. Soong
中科院分区:
--
文献类型:
--
作者:
J. Chen;F. Soong

文献摘要

被引文献

相似文献

作者提出了一种基于N最佳候选的判别训练方法,用于构造高性能的HMM语音识别器。该算法有两个明显的特点:N-最好的假设用于训练判别模型;和一个新的帧级损失函数最小化,以提高正确和不正确的假设之间的分离。最好的N个候选解码的基础上,他们最近提出的树格快速搜索算法。新的帧级损失函数被定义为正确假设和竞争假设之间的半波整流对数似然差,在所有训练令牌上被最小化。通过沿梯度下降方向沿着调整HMM参数来实现最小化。两个语音识别应用程序已经过测试,包括一个独立于说话人的,小词汇量(10个普通话数字),连续语音识别,和一个说话人训练,大词汇量(5000常用的中文单词),孤立词识别。与传统的最大似然训练HISTORY相比,性能得到了显着提高。在汉语连续数字识别实验中,未知长度解码的字符串错误率从17.0%降到10.8%,已知长度解码的字符串错误率从8.2%降到5.2%。在大词汇量孤立词识别实验中,识别错误率从7.2%降低到3.8%。此外,他们发现,在准备N-最佳假设时使用更宽松的解码约束会产生更好的识别结果。>
The authors propose an N-best candidates-based discriminative training procedure for constructing high-performance HMM speech recognizers. The algorithm has two distinct features: N-best hypotheses are used for training discriminative models; and a new frame-level loss function is minimized to improve the separation between the correct and incorrect hypotheses. The N-best candidates are decoded based on their recently proposed tree-trellis fast search algorithm. The new frame-level loss function, which is defined as a halfwave rectified log-likelihood difference between the correct and competing hypotheses, is minimized over all training tokens. The minimization is carried out by adjusting the HMM parameters along a gradient descent direction. Two speech recognition applications have been tested, including a speaker independent, small vocabulary (ten Mandarin Chinese digits), continuous speech recognition, and a speaker-trained, large vocabulary (5000 commonly used Chinese words), isolated word recognition. Significant performance improvement over the traditional maximum likelihood trained HMMs has been obtained. In the connected Chinese digit recognition experiment, the string error rate is reduced from 17.0 to 10.8% for unknown length decoding and from 8.2 to 5.2% for known length decoding. In the large vocabulary, isolated word recognition experiment, the recognition error rate is reduced from 7.2 to 3.8%. Additionally, they have found that using more relaxed decoding constraints in preparing N-best hypotheses yields better recognition results. >