CNN-based virtual microphone signal estimation for MPDR beamforming in underdetermined situations

CNN-based virtual microphone signal estimation for MPDR beamforming in underdetermined situations
复制标题

DOI:
10.23919/eusipco.2019.8903040
复制
发表时间:
2019-09
期刊:
2019 27th European Signal Processing Conference (EUSIPCO)
影响因子:
--
通讯作者:
K. Yamaoka;Li Li-Li;Nobutaka Ono;S. Makino;Takeshi Yamada
K. Yamaoka;Li Li-Li;Nobutaka Ono;S. Makino;Takeshi Yamada
中科院分区:
其他
文献类型:
--
作者:
K. Yamaoka;Li Li-Li;Nobutaka Ono;S. Makino;Takeshi Yamada

文献摘要

相似文献

在本文中,我们提出了一种新的方法,实际上增加了两个真实麦克风之间的麦克风元素的数量,以提高在欠确定情况下的语音增强性能。虚拟麦克风技术是最近提出的一种基于β散度的虚拟麦克风技术,它通过线性内插相位和非线性内插幅度来估计音频信号域中的虚拟信号,实验表明该技术在改善语音增强性能方面是有效的。此外,有报道称,随着非线性的改善,性能也趋于改善。然而,这种方法的一个缺点是在每个时频区间独立地进行内插,忽略了语音信号的谱结构和时间结构。为了解决这一问题,并改善非线性,受神经网络对非线性函数和语音谱图建模的强大能力的启发,本文提出了一种替代的幅度内插方法。在该方法中,我们使用卷积神经网络作为幅度估计器,最小化最小功率无失真响应(MPDR)波束形成器的输出与目标语音信号之间的均方误差。实验结果表明,该方法在提高语音增强性能方面具有很大的潜力,不仅优于传统的虚拟麦克风技术,而且在相应的确定情况下的性能也优于传统的虚拟麦克风技术。
In this paper, we propose a novel approach to virtually increasing the number of microphone elements between two real microphones to improve speech enhancement performance in underdetermined situations. The virtual microphone technique, with which virtual signals in the audio signal domain are estimated by linearly interpolating the phase and nonlinearly interpolating the amplitude independently on the basis of β-divergence, has been recently proposed and experimentally shown to be effective in improving speech enhancement performance. Furthermore, it has been reported that the performance tends to improve as the nonlinearity is improved. However, one drawback of this method is that the interpolation is employed in each time-frequency bin independently, in which the spectral and temporal structures of speech signals are ignored. To address this problem and improve the nonlinearity, motivated by the high capability of neural networks to model nonlinear functions and speech spectrograms, in this paper, we propose an alternative method of amplitude interpolation. In this method, we employ a convolutional neural network as an amplitude estimator that minimizes the mean squared error between the outputs of a minimum power distortionless response (MPDR) beamformer and the target speech signals. The experimental results revealed that the proposed method showed high potential for improving speech enhancement performance, which was not only superior to that of the conventional virtual microphone technique but also the performance in the corresponding determined situation.