Adaptive Very Deep Convolutional Residual Network for Noise Robust Speech Recognition

Adaptive Very Deep Convolutional Residual Network for Noise Robust Speech Recognition
复制标题

用于噪声鲁棒语音识别的自适应超深卷积残差网络

DOI:
10.1109/taslp.2018.2825432
复制
发表时间:
2018-08
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Kai Yu
Kai Yu
中科院分区:
其他
文献类型:
--
作者:
Tian Tan;Yanmin Qian;Hu Hu;Ying Zhou;Wen Ding;Kai Yu

文献摘要

参考文献

被引文献

相似文献

虽然语音识别技术已经取得了很大的进步,但在噪声环境下,语音识别的性能仍然存在很大的下降。我们以前的工作已经证明了非常深的卷积神经网络(VDCNN)的上级噪声鲁棒性。基于我们在VDCNN上的工作,本文提出了一种更先进的模型,称为非常深的卷积残差网络(VDCRN)。该模型结合了批量归一化和残差学习,比以往的VDCNNs具有更强的鲁棒性。然后,为了缓解训练和测试条件之间的不匹配,本文对新的VDCRN进行了模型自适应和自适应训练,并进行了比较。本文重点研究了因子感知训练(FAT)和聚类自适应训练(CAT)。对于FAT,探索了一个统一的框架。对于CAT,两个计划首先探讨构建基地的规范模型,此外,CAT的分解版本的设计,以解决多个非语音变量在一个模型。最后,提出了一个完整的多通道系统,以实现最佳的系统性能在嘈杂的情况下。在Aurora4(具有加性噪声和信道失真的模拟数据)、CHiME4(具有加性噪声和混响的模拟数据和真实的数据)以及AMI会议转录任务(具有显著混响的真实的数据)上对所提出的新方法进行了评估,评估不仅包括不同的噪声条件,而且包括模拟和真实的噪声数据。实验结果表明,新的VDCRN是更强大的,在此模型上的自适应可以进一步显着降低字错误率(WER)。与基线VDCNN或长短期记忆相比,所提出的最佳架构在所有任务上都获得了一致且非常大的改进。特别是在Aurora 4上,仅通过改进声学建模就实现了新的里程碑5.67% WER。
Although great progress has been made in automatic speech recognition, significant performance degradation still exists in noisy environments. Our previous work has demonstrated the superior noise robustness of very deep convolutional neural networks (VDCNN). Based on our work on VDCNNs, this paper proposes a more advanced model referred to as the very deep convolutional residual network (VDCRN). This new model incorporates batch normalization and residual learning, showing more robustness than previous VDCNNs.Then, to alleviate the mismatch between the training and testing conditions, model adaptation and adaptive training are developed and compared for the new VDCRN. This paper focuses on factor aware training (FAT) and cluster adaptive training (CAT). For FAT, a unified framework is explored. For CAT, two schemes are first explored to construct the bases in the canonical model; furthermore, a factorized version of CAT is designed to address multiple nonspeech variabilities in one model. Finally, a complete multipass system is proposed to achieve the best system performance in the noisy scenarios. The proposed new approaches are evaluated on three different tasks: Aurora4 (simulated data with additive noise and channel distortion), CHiME4 (both simulated and real data with additive noise and reverberation), and the AMI meeting transcription task (real data with significant reverberation).The evaluation not only includes different noisy conditions, but also covers both simulated and real noisy data. The experiments show that the new VDCRN is more robust, and the adaptation on this model can further significantly reduce the word error rate (WER). The proposed best architecture obtains consistent and very large improvements on all tasks compared to the baseline VDCNN or long short-term memory. Particularly, on Aurora4 a new milestone 5.67% WER is achieved by only improving acoustic modeling.
DOI: 10.1109/icassp.2015.7178787
发表时间: 2015-04
期刊: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
Tian Tan;Y. Qian;Maofan Yin;Yimeng Zhuang;Kai Yu
通讯作者: Tian Tan;Y. Qian;Maofan Yin;Yimeng Zhuang;Kai Yu
基于神经网络的多因素感知联合训练,实现鲁棒语音识别
DOI: 10.1109/taslp.2016.2598308
发表时间: 2016-12
影响因子: 5.4
作者:
Yanmin Qian;Tian Tan;Dong Yu
通讯作者: Dong Yu
DOI: 10.1109/slt.2012.6424251
发表时间: 2012-12
期刊: 2012 IEEE Spoken Language Technology Workshop (SLT)
影响因子: --
作者:
K. Yao;Dong Yu;F. Seide;Hang Su;L. Deng;Y. Gong
通讯作者: K. Yao;Dong Yu;F. Seide;Hang Su;L. Deng;Y. Gong
DOI: 10.1109/icassp.2008.4518541
发表时间: 2008-05
期刊: 2008 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子: --
作者:
Dong Yu;L. Deng;J. Droppo;Jian Wu;Y. Gong;A. Acero
通讯作者: Dong Yu;L. Deng;J. Droppo;Jian Wu;Y. Gong;A. Acero
DOI: 10.1109/tasl.2012.2198059
发表时间: 2012-09
期刊: IEEE Transactions on Audio, Speech, and Language Processing
影响因子: --
作者:
Yongqiang Wang;M. Gales
通讯作者: Yongqiang Wang;M. Gales