Adaptive Very Deep Convolutional Residual Network for Noise Robust Speech Recognition
Adaptive Very Deep Convolutional Residual Network for Noise Robust Speech Recognition
复制标题
用于噪声鲁棒语音识别的自适应超深卷积残差网络
DOI:
10.1109/taslp.2018.2825432
复制
发表时间:
2018-08
期刊:
影响因子:
--
通讯作者:
Kai Yu
中科院分区:
文献类型:
--
作者:
Tian Tan;Yanmin Qian;Hu Hu;Ying Zhou;Wen Ding;Kai Yu
Although great progress has been made in automatic speech recognition, significant performance degradation still exists in noisy environments. Our previous work has demonstrated the superior noise robustness of very deep convolutional neural networks (VDCNN). Based on our work on VDCNNs, this paper proposes a more advanced model referred to as the very deep convolutional residual network (VDCRN). This new model incorporates batch normalization and residual learning, showing more robustness than previous VDCNNs.Then, to alleviate the mismatch between the training and testing conditions, model adaptation and adaptive training are developed and compared for the new VDCRN. This paper focuses on factor aware training (FAT) and cluster adaptive training (CAT). For FAT, a unified framework is explored. For CAT, two schemes are first explored to construct the bases in the canonical model; furthermore, a factorized version of CAT is designed to address multiple nonspeech variabilities in one model. Finally, a complete multipass system is proposed to achieve the best system performance in the noisy scenarios. The proposed new approaches are evaluated on three different tasks: Aurora4 (simulated data with additive noise and channel distortion), CHiME4 (both simulated and real data with additive noise and reverberation), and the AMI meeting transcription task (real data with significant reverberation).The evaluation not only includes different noisy conditions, but also covers both simulated and real noisy data. The experiments show that the new VDCRN is more robust, and the adaptation on this model can further significantly reduce the word error rate (WER). The proposed best architecture obtains consistent and very large improvements on all tasks compared to the baseline VDCNN or long short-term memory. Particularly, on Aurora4 a new milestone 5.67% WER is achieved by only improving acoustic modeling.
登录
查看更多内容
DOI:
10.1109/icassp.2015.7178787
发表时间:
2015-04
期刊:
2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
Tian Tan;Y. Qian;Maofan Yin;Yimeng Zhuang;Kai Yu
通讯作者:
Tian Tan;Y. Qian;Maofan Yin;Yimeng Zhuang;Kai Yu
DOI:
10.1109/taslp.2016.2598308
发表时间:
2016-12
影响因子:
5.4
作者:
Yanmin Qian;Tian Tan;Dong Yu
通讯作者:
Dong Yu
DOI:
10.1109/slt.2012.6424251
发表时间:
2012-12
期刊:
2012 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
作者:
K. Yao;Dong Yu;F. Seide;Hang Su;L. Deng;Y. Gong
通讯作者:
K. Yao;Dong Yu;F. Seide;Hang Su;L. Deng;Y. Gong
DOI:
10.1109/icassp.2008.4518541
发表时间:
2008-05
期刊:
2008 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
作者:
Dong Yu;L. Deng;J. Droppo;Jian Wu;Y. Gong;A. Acero
通讯作者:
Dong Yu;L. Deng;J. Droppo;Jian Wu;Y. Gong;A. Acero
DOI:
10.1109/tasl.2012.2198059
发表时间:
2012-09
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
Yongqiang Wang;M. Gales
通讯作者:
Yongqiang Wang;M. Gales