Active and unsupervised learning for automatic speech recognition

Active and unsupervised learning for automatic speech recognition
复制标题

DOI:
10.21437/eurospeech.2003-552
复制
发表时间:
2003-09
期刊:
2003 IEEE Workshop on Automatic Speech Recognition and Understanding (IEEE Cat. No.03EX721)
影响因子:
--
通讯作者:
G. Riccardi;Dilek Z. Hakkani-Tür
G. Riccardi;Dilek Z. Hakkani-Tür
中科院分区:
其他
文献类型:
--
作者:
G. Riccardi;Dilek Z. Hakkani-Tür

文献摘要

被引文献

相似文献

最先进的语音识别系统是使用语音话语的人类翻译来训练的。在本文中,我们描述了一种方法,结合联合收割机主动和无监督学习的自动语音识别(ASR)。我们的目标是最大限度地减少训练声学和语言模型的人类监督,并最大限度地提高转录和未转录数据的性能。主动学习的目的是通过自动处理未标记的示例来减少要标记的训练示例的数量,然后选择相对于给定成本函数信息量最大的示例。对于无监督学习,我们通过使用其ASR输出和单词置信度得分来利用剩余的未转录数据。我们的实验表明,通过结合主动学习和无监督学习,给定单词准确度所需的标记数据量可以减少75%。
State-of-the-art speech recognition systems are trained us-ing human transcriptions of speech utterances. In this paper, we describe a method to combine active and unsupervised learning for automatic speech recognition (ASR). The goal is to minimize the human supervision for training acoustic and language models and to maximize the performance given the transcribed and untranscribed data. Active learning aims at reducing the number of training examples to be labeled by automatically processing the unlabeled examples, and then selecting the most informative ones with respect to a given cost function. For unsupervised learning, we utilize the remaining untranscribed data by using their ASR output and word confidence scores. Our experiments show that the amount of labeled data needed for a given word accuracy can be reduced by 75% by combining active and unsupervised learning.