A Probabilistic Model for Robust Localization Based on a Binaural Auditory Front-End

A Probabilistic Model for Robust Localization Based on a Binaural Auditory Front-End
复制标题

DOI:
10.1109/tasl.2010.2042128
复制
发表时间:
2011
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
T. May;S. Par;A. Kohlrausch
T. May;S. Par;A. Kohlrausch
中科院分区:
其他
文献类型:
--
作者:
T. May;S. Par;A. Kohlrausch

文献摘要

被引文献

相似文献

虽然在基于机器的定位领域已经进行了广泛的研究,混响和多个源的存在对定位性能的降级效果仍然是一个主要问题。受人类听觉系统对复杂声学场景进行鲁棒分析的能力的启发,本文使用相关的外围级作为基于双耳信号的声源方位角估计的前端。在水平面中定位声源的一种经典方法是通过搜索互相关函数中的最大值来估计双耳之间的耳间时间差(ITD)。除了ITD之外,耳间水平差(ILD)也可能有助于定位,特别是在波长变得小于头部直径的较高频率下,导致模糊的ITD信息。ITD和ILD对方位角的相互依赖性是一种复杂的模式,也取决于室内声学,因此可以通过方位角相关高斯混合模型(GARCH)来学习。进行多条件训练以考虑由多个源产生的双耳特征的可变性和混响的影响。所提出的定位模型优于国家的最先进的定位技术在模拟不利的声学条件。
Although extensive research has been done in the field of machine-based localization, the degrading effect of reverberation and the presence of multiple sources on localization performance has remained a major problem. Motivated by the ability of the human auditory system to robustly analyze complex acoustic scenes, the associated peripheral stage is used in this paper as a front-end to estimate the azimuth of sound sources based on binaural signals. One classical approach to localize an acoustic source in the horizontal plane is to estimate the interaural time difference (ITD) between both ears by searching for the maximum in the cross-correlation function. Apart from ITDs, the interaural level difference (ILD) can contribute to localization, especially at higher frequencies where the wavelength becomes smaller than the diameter of the head, leading to ambiguous ITD information. The interdependency of ITD and ILD on azimuth is a complex pattern that depends also on the room acoustics, and is therefore learned by azimuth-dependent Gaussian mixture models (GMMs). Multiconditional training is performed to take into account the variability of the binaural features which results from multiple sources and the effect of reverberation. The proposed localization model outperforms state-of-the-art localization techniques in simulated adverse acoustic conditions.