Unsupervised training of acoustic models for large vocabulary continuous speech recognition

Unsupervised training of acoustic models for large vocabulary continuous speech recognition
复制标题

DOI:
10.1109/asru.2001.1034648
复制
发表时间:
2001-12
期刊:
IEEE Workshop on Automatic Speech Recognition and Understanding, 2001. ASRU '01.
影响因子:
--
通讯作者:
F. Wessel;H. Ney
F. Wessel;H. Ney
中科院分区:
其他
文献类型:
--
作者:
F. Wessel;H. Ney

文献摘要

被引文献

相似文献

对于语音识别系统,声学训练数据的量是至关重要的。在过去,大量的语音被记录和人工转录用于训练。针对目前存在的各种形式的非转录语音,研究了基于已识别转录的语音识别器的无监督训练问题。一个低成本的识别器训练,只有一个小时的人工转录的语音被用来识别72小时的未转录的声学数据。然后,这些transnumbers与置信度措施相结合,训练一个改进的识别器。系统地研究了用于检测可能的识别错误的置信度的作用。最后,迭代地应用无监督训练。使用这种方法,识别器的训练,很少的人工努力,而损失只有14.3%的相对广播新闻'96和18.6%的相对广播新闻'98评估测试集。
For speech recognition systems, the amount of acoustic training data is of crucial importance. In the past, large amounts of speech were recorded and transcribed manually for training. Since untranscribed speech is available in various forms these days, the unsupervised training of a speech recognizer on recognized transcriptions is studied. A low-cost recognizer trained with only one hour of manually transcribed speech is used to recognize 72 hours of untranscribed acoustic data. These transcriptions are then used in combination with confidence measures to train an improved recognizer. The effect of confidence measures which are used to detect possible recognition errors is studied systematically. Finally, the unsupervised training is applied iteratively. Using this method, the recognizer is trained with very little manual effort while losing only 14.3% relative on the Broadcast News '96 and 18.6% relative on the Broadcast News '98 evaluation test sets.