A generic neural acoustic beamforming architecture for robust multi-channel speech processing

A generic neural acoustic beamforming architecture for robust multi-channel speech processing
复制标题

DOI:
10.1016/j.csl.2016.11.007
复制
发表时间:
2017-11-01
影响因子:
4.3
通讯作者:
Haeb-Umbach, Reinhold
Haeb-Umbach, Reinhold
中科院分区:
计算机科学3区
文献类型:
--
作者:
Heymann, Jahn;Drude, Lukas;Haeb-Umbach, Reinhold

文献摘要

被引文献

相似文献

声波束形成可以在多通道情况下大大提高自动语音识别(ASR)和语音增强系统的性能。我们最近提出了一种方法来支持基于模型的广义特征值波束形成操作与一个强大的神经网络的频谱掩模估计。增强系统具有许多期望的特性。特别地,不需要对声学传递函数的性质进行任何假设(例如,是无回声的),也不需要知道阵列配置。虽然该系统最初是为了在嘈杂的环境中增强语音而开发的,但我们在本文中表明,它在抑制混响方面也是有效的,从而导致了一个通用的可训练多通道语音增强系统,用于鲁棒的语音处理。为了支持这一说法,我们考虑了两个不同的数据集:CHiME 3challenge,它具有挑战性的现实世界的噪声失真,以及Reverbchallenge,它专注于混响引起的失真。我们评估系统的语音增强和识别任务。对于第一个任务,我们提出了一种新的方法来科普由广义特征值波束形成器引入的失真,通过重新归一化每个频率点的目标能量,并测量其有效性的PESQ得分。对于后者,我们将增强的信号馈送到强大的DNN后端,并在两个数据集上实现最先进的ASR结果。我们进一步实验了不同的网络架构进行频谱掩模估计:一个只有一个隐藏层的小型前馈网络,一个卷积神经网络和一个双向长短期记忆网络,表明即使是一个小型网络也能够提供显着的性能改进。(C)2017爱思唯尔有限公司版权所有
Acoustic beamforming can greatly improve the performance of Automatic Speech Recognition(ASR) and speech enhancement systems when multiple channels are available. We recently proposed a way to support the model-based Generalized Eigenvalue beamforming operation with a powerful neural network for spectral mask estimation. The enhancement system has a number of desirable properties. In particular, neither assumptions need to be made about the nature of the acoustic transfer function (e.g., being anechonic), nor does the array configuration need to be known. While the system has been originally developed to enhance speech in noisy environments, we show in this article that it is also effective in suppressing reverberation, thus leading to a generic trainable multi-channel speech enhancement system for robust speech processing. To support this claim, we consider two distinct datasets: The CHiME 3challenge, which features challenging real-world noise distortions, and the Reverbchallenge, which focuses on distortions caused by reverberation. We evaluate the system both with respect to a speech enhancement and a recognition task. For the first task we propose a new way to cope with the distortions introduced by the Generalized Eigenvalue beamformer by renormalizing the target energy for each frequency bin, and measure its effectiveness in terms of the PESQ score. For the latter we feed the enhanced signal to a strong DNN back-end and achieve state-of-the-art ASR results on both datasets. We further experiment with different network architectures for spectral mask estimation: One small feed-forward network with only one hidden layer, one Convolutional Neural Network and one bi-directional Long Short-Term Memory network, showing that even a small network is capable of delivering significant performance improvements. (C) 2017 Elsevier Ltd. All rights reserved.