Optimizing neural-network supported acoustic beamforming by algorithmic differentiation

Optimizing neural-network supported acoustic beamforming by algorithmic differentiation
复制标题

DOI:
10.1109/icassp.2017.7952140
复制
发表时间:
2017-03
期刊:
2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Christoph Boeddeker;Patrick Hanebrink;Lukas Drude;Jahn Heymann;Reinhold Häb-Umbach
Christoph Boeddeker;Patrick Hanebrink;Lukas Drude;Jahn Heymann;Reinhold Häb-Umbach
中科院分区:
其他
文献类型:
--
作者:
Christoph Boeddeker;Patrick Hanebrink;Lukas Drude;Jahn Heymann;Reinhold Häb-Umbach

文献摘要

被引文献

相似文献

在本文中,我们展示了如何通过算法微分来优化用于声波束形成器频谱模板估计的神经网络。使用波束形成器输出 SNR 作为最大化的目标函数,梯度通过波束形成器一直传播到神经网络,神经网络提供干净的语音和噪声掩模,通过特征值分解来估计波束形成器系数。一个关键的理论结果是涉及复值特征向量的特征值问题的导数。 CHiME-3挑战数据库上的实验结果证明了该方法的有效性。本文开发的工具是语音增强和语音识别端到端优化的关键组成部分。
In this paper we show how a neural network for spectral mask estimation for an acoustic beamformer can be optimized by algorithmic differentiation. Using the beamformer output SNR as the objective function to maximize, the gradient is propagated through the beamformer all the way to the neural network which provides the clean speech and noise masks from which the beamformer coefficients are estimated by eigenvalue decomposition. A key theoretical result is the derivative of an eigenvalue problem involving complex-valued eigenvectors. Experimental results on the CHiME-3 challenge database demonstrate the effectiveness of the approach. The tools developed in this paper are a key component for an end-to-end optimization of speech enhancement and speech recognition.