Speaking rate compensation based on likelihood criterion in acoustic model training and decoding

Speaking rate compensation based on likelihood criterion in acoustic model training and decoding
复制标题

声学模型训练和解码中基于似然准则的语速补偿

DOI:
10.21437/icslp.2002-343
复制
发表时间:
2002
期刊:
--
影响因子:
--
通讯作者:
Satoshi Nakamura
Satoshi Nakamura
中科院分区:
--
文献类型:
--
作者:
K. Okuda;Tatsuya Kawahara;Satoshi Nakamura

文献摘要

被引文献

相似文献

在本文中,我们提出了一种使用帧周期和帧长自适应的语速补偿方法。我们的方法解码的输入话语使用几组帧周期和帧长度参数的语音分析。然后,该方法选择具有最高得分的最佳集合,该集合由经帧周期归一化的声学似然、语言似然和插入惩罚组成。此外,我们将这种方法应用于声学模型的训练。我们使用Viterbi对齐计算每个帧周期和帧长度的声学似然,并为每个训练话语选择最佳的一个。建议的语速补偿应用到声学模型创建过程和解码过程中,导致自发演讲语音识别任务的准确率提高了2.9%(绝对)。
In this paper, we propose a speaking rate compensation method using frame period and frame length adaptation. Our method decodes an input utterance using several sets of frame period and frame length parameters for speech analysis. Then, this method selects the best set with the highest score which consists of the acoustic likelihood normalized by frame period, language likelihood and insertion penalty. Furthermore, we apply this approach to the training of the acoustic model. We calculate the acoustic likelihood for each frame period and frame length using Viterbi alignment and select the best one for each training utter-ance. The proposed speaking rate compensation applied to both the acoustic model creation process and decoding process resulted in accuracy improvement of 2.9% (absolute) for spontaneous lecture speech recognition task.