Adaptation of context-dependent deep neural networks for automatic speech recognition

Adaptation of context-dependent deep neural networks for automatic speech recognition
复制标题

DOI:
10.1109/slt.2012.6424251
复制
发表时间:
2012-12
期刊:
2012 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
K. Yao;Dong Yu;F. Seide;Hang Su;L. Deng;Y. Gong
K. Yao;Dong Yu;F. Seide;Hang Su;L. Deng;Y. Gong
中科院分区:
其他
文献类型:
--
作者:
K. Yao;Dong Yu;F. Seide;Hang Su;L. Deng;Y. Gong

文献摘要

被引文献

相似文献

本文评估了上下文相关深度神经网络隐马尔可夫模型(CD-DNN-HMM)的自适应方法在自动语音识别中的有效性。我们研究了仿射变换及其几种变体以适应顶层隐藏层。我们将仿射变换与Softmax层权重的直接自适应进行了比较。对特征空间判别线性回归(FDLR)方法的输入层仿射变换进行了评价。在一个大词汇量的语音识别任务中,FDLR和顶层隐含层自适应的随机梯度上升实现分别比基准DNN性能降低了17%和14%的错字率(WERS)。通过批量更新实施,Softmax层适配技术减少了10%的WER。我们观察到,使用偏置移位的效果与缩放加偏置移位的效果一样好。
In this paper, we evaluate the effectiveness of adaptation methods for context-dependent deep-neural-network hidden Markov models (CD-DNN-HMMs) for automatic speech recognition. We investigate the affine transformation and several of its variants for adapting the top hidden layer. We compare the affine transformations against direct adaptation of the softmax layer weights. The feature-space discriminative linear regression (fDLR) method with the affine transformations on the input layer is also evaluated. On a large vocabulary speech recognition task, a stochastic gradient ascent implementation of the fDLR and the top hidden layer adaptation is shown to reduce word error rates (WERs) by 17% and 14%, respectively, compared to the baseline DNN performances. With a batch update implementation, the softmax layer adaptation technique reduces WERs by 10%. We observe that using bias shift performs as well as doing scaling plus bias shift.