Unsupervised training of acoustic models for large vocabulary continuous speech recognition
Unsupervised training of acoustic models for large vocabulary continuous speech recognition
复制标题
DOI:
10.1109/asru.2001.1034648
复制
发表时间:
2001-12
期刊:
影响因子:
--
通讯作者:
F. Wessel;H. Ney
中科院分区:
文献类型:
--
作者:
F. Wessel;H. Ney
For speech recognition systems, the amount of acoustic training data is of crucial importance. In the past, large amounts of speech were recorded and transcribed manually for training. Since untranscribed speech is available in various forms these days, the unsupervised training of a speech recognizer on recognized transcriptions is studied. A low-cost recognizer trained with only one hour of manually transcribed speech is used to recognize 72 hours of untranscribed acoustic data. These transcriptions are then used in combination with confidence measures to train an improved recognizer. The effect of confidence measures which are used to detect possible recognition errors is studied systematically. Finally, the unsupervised training is applied iteratively. Using this method, the recognizer is trained with very little manual effort while losing only 14.3% relative on the Broadcast News '96 and 18.6% relative on the Broadcast News '98 evaluation test sets.