Robust Binaural Localization of a Target Sound Source by Combining Spectral Source Models and Deep Neural Networks

Robust Binaural Localization of a Target Sound Source by Combining Spectral Source Models and Deep Neural Networks
复制标题

DOI:
10.1109/taslp.2018.2855960
复制
发表时间:
2018-11-01
影响因子:
5.4
通讯作者:
Brown, Guy J.
Brown, Guy J.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ma, Ning;Gonzalez, Jose A.;Brown, Guy J.

文献摘要

被引文献

相似文献

尽管有明确的证据表明生物空间听力中存在自上而下(例如注意力)效应,但相对较少的机器听力系统利用声音定位中自上而下的基于模型的知识。本文通过提出一种新颖的双耳声音定位框架来解决这个问题,该框架结合了有关声源频谱特征的基于模型的信息和深度神经网络(DNN)。首先在训练阶段使用从孤立的声音信号中提取的频谱特征来估计目标源模型和背景源模型。当背景源的身份不可用时,可以使用通用背景模型。在测试过程中,联合使用源模型来解释混合观测结果,并通过选择性地加权基于 DNN 的定位系统输出的源方位角后验来改进定位过程。为了解决训练和测试之间可能存在的不匹配问题,在测试期间进一步采用动态模型适应过程,以迭代方式直接从噪声观测中调整背景模型参数。因此,所提出的系统将基于模型和数据驱动的信息流结合在单个计算框架内。评估任务涉及在存在干扰源和房间混响的情况下定位目标语音源。我们的实验表明,通过以这种方式利用基于模型的信息,可以在各种噪声和混响条件下显着提高声音定位性能。
Despite there being a clear evidence for top-down (e.g., attentional) effects in biological spatial hearing, relatively few machine hearing systems exploit the top-down model-based knowledge in sound localization. This paper addresses this issue by proposing a novel framework for the binaural sound localization that combines the model-based information about the spectral characteristics of sound sources and deep neural networks (DNNs). A target source model and a background source model are first estimated during a training phase using spectral features extracted from sound signals in isolation. When the identity of the background source is not available, a universal background model can be used. During testing, the source models are used jointly to explain the mixed observations and improve the localization process by selectively weighting source azimuth posteriors output by a DNN-based localization system. To address the possible mismatch between the training and testing, a model adaptation process is further employed the on-the-fly during testing, which adapts the background model parameters directly from the noisy observations in an iterative manner. The proposed system, therefore, combines the model-based and data-driven information flow within a single computational framework. The evaluation task involved localization of a target speech source in the presence of an interfering source and room reverberation. Our experiments show that by exploiting the model-based information in this way, the sound localization performance can be improved substantially under various noisy and reverberant conditions.