Robust Speech Recognition Using Generalized Distillation Framework

Robust Speech Recognition Using Generalized Distillation Framework
复制标题

DOI:
10.21437/interspeech.2016-852
复制
发表时间:
2016-09
期刊:
2023 IEEE International Symposium on Circuits and Systems (ISCAS)
影响因子:
--
通讯作者:
K. Markov;T. Matsui
K. Markov;T. Matsui
中科院分区:
其他
文献类型:
--
作者:
K. Markov;T. Matsui

文献摘要

被引文献

相似文献

本文提出了一种基于广义蒸馏框架的噪声鲁棒语音识别系统。假设在训练过程中,除了训练数据之外,还有某种“特权”信息可用,可以用来指导训练过程。这允许获得在测试时优于仅基于常规训练数据构建的系统的系统。在有噪声的语音识别任务中,特权信息是从一个被称为“教师”的模型中获得的,该模型只对干净的语音进行训练。规则模型称为“学生”,它在嘈杂的话语上进行训练,并将教师的输出用于相应的干净话语。因此,对于该框架,需要并行的干净/噪声语音数据。我们在提供此类数据的Aurora2数据库上进行了实验。我们的系统使用混合DNN-HMM声学模型,其中神经网络在解码过程中提供HMM状态概率。教师DNN是在干净的数据上训练的,而学生DNN是使用多条件(各种信噪比)数据训练的。学生DNN损失函数结合了从训练数据的强制对齐中获得的目标和教师DNN在输入相应的干净特征时的输出。实验结果清楚地表明,蒸馏框架是有效的,可以显著降低单词错误率。
In this paper, we propose a noise robust speech recognition system built using generalized distillation framework. It is assumed that during training, in addition to the training data, some kind of ”privileged” information is available and can be used to guide the training process. This allows to obtain a system which at test time outperforms those built on regular training data alone. In the case of noisy speech recognition task, the privileged information is obtained from a model, called ”teacher”, trained on clean speech only. The regular model, called ”student”, is trained on noisy utterances and uses teacher’s output for the corresponding clean utterances. Thus, for this framework a parallel clean/noisy speech data are required. We experimented on the Aurora2 database which provides such kind of data. Our system uses hybrid DNN-HMM acoustic model where neural networks provide HMM state probabilities during decoding. The teacher DNN is trained on the clean data, while the student DNN is trained using multi-condition (various SNRs) data. The student DNN loss function combines the targets obtained from forced alignment of the training data and the outputs of the teacher DNN when fed with the corresponding clean features. Experimental results clearly show that distillation framework is effective and allows to achieve significant reduction in the word error rate.