Robust speech recognition using beamforming with adaptive microphone gains and multichannel noise reduction

Robust speech recognition using beamforming with adaptive microphone gains and multichannel noise reduction
复制标题

DOI:
10.1109/asru.2015.7404831
复制
发表时间:
2015-12
期刊:
2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU)
影响因子:
--
通讯作者:
Shengkui Zhao;Xiong Xiao;Zhaofeng Zhang;Thi Ngoc Tho Nguyen;X. Zhong;Bo Ren;Longbiao Wang;Douglas L. Jones;Chng Eng Siong;Haizhou Li
Shengkui Zhao;Xiong Xiao;Zhaofeng Zhang;Thi Ngoc Tho Nguyen;X. Zhong;Bo Ren;Longbiao Wang;Douglas L. Jones;Chng Eng Siong;Haizhou Li
中科院分区:
其他
文献类型:
--
作者:
Shengkui Zhao;Xiong Xiao;Zhaofeng Zhang;Thi Ngoc Tho Nguyen;X. Zhong;Bo Ren;Longbiao Wang;Douglas L. Jones;Chng Eng Siong;Haizhou Li

文献摘要

相似文献

本文针对第三届 CHiME 挑战赛提出了一种使用麦克风阵列的强大语音识别系统。提出了一种具有自适应麦克风增益的最小方差无失真响应(MVDR)波束形成器,以实现鲁棒的波束形成。使用语音主导的时频仓研究了两种麦克风增益估计方法。还提出了多通道降噪(MCNR)后处理,以进一步减少 MVDR 处理信号中的干扰。 ChiME-3 挑战赛的实验结果表明,所提出的具有麦克风增益的 MVDR 波束形成器和 MCNR 后处理都显着提高了语音识别性能。凭借最先进的基于深度神经网络 (DNN) 的声学模型,我们的系统在评估集的真实测试数据上实现了 11.67% 的单词错误率 (WER)。
This paper presents a robust speech recognition system using a microphone array for the 3rd CHiME Challenge. A minimum variance distortionless response (MVDR) beamformer with adaptive microphone gains is proposed for robust beamforming. Two microphone gain estimation methods are studied using the speech-dominant time-frequency bins. A multichannel noise reduction (MCNR) postprocessing is also proposed to further reduce the interference in the MVDR processed signal. Experimental results for the ChiME-3 challenge show that both the proposed MVDR beamformer with microphone gains and the MCNR postprocessing improve the speech recognition performance significantly. With the state-of-the-art deep neural network (DNN) based acoustic model, our system achieves a word error rate (WER) of 11.67% on the real test data of the evaluation set.