Deep learning assisted sound source localization using two orthogonal first-order differential microphone arrays.

Deep learning assisted sound source localization using two orthogonal first-order differential microphone arrays.
复制标题

使用两个正交一阶差分麦克风阵列进行深度学习辅助声源定位。

DOI:
10.1121/10.0003445
复制
发表时间:
2021
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
通讯作者:
Yanwen Li
Yanwen Li
中科院分区:
--
文献类型:
--
作者:
Nian Liu;Huawei Chen;Kunkun SongGong;Yanwen Li

文献摘要

被引文献

相似文献

声源定位在嘈杂和混响的房间使用麦克风阵列仍然是一个具有挑战性的任务,特别是对于小尺寸的阵列。近年来,通过将声音定位问题重新定义为分类问题,深度学习辅助方法取得了可喜的进展。基于深度学习的方法的关键在于在噪声和混响条件下有效地提取声音位置特征。普遍采用的功能是基于公认的广义互相关相位变换(GCC-PHAT),这是已知的是有助于打击房间混响。然而,GCC-PHAT功能可能不适用于小型阵列。本文提出了一种基于深度学习的声音定位方法,该方法使用由两个正交的一阶差分麦克风阵列构成的小尺寸麦克风阵列。提出了一种改进的基于声强估计的特征提取方案,通过在白化加权构造中解耦声压分量和质点振速分量之间的相关性,增强了时频区间声强特征的鲁棒性。仿真和真实世界的实验结果表明,所提出的深度学习辅助方法可以实现更高的空间分辨率,并且上级在噪声和混响环境中使用GCC-PHAT或小尺寸阵列的声音强度特征的最先进的方法。
Sound source localization in noisy and reverberant rooms using microphone arrays remains a challenging task, especially for small-sized arrays. Recent years have seen promising advances on deep learning assisted approaches by reformulating the sound localization problem as a classification one. A key to the deep learning-based approaches lies in extracting sound location features effectively in noisy and reverberant conditions. The popularly adopted features are based on the well-established generalized cross correlation phase transform (GCC-PHAT), which is known to be helpful in combating room reverberation. However, the GCC-PHAT features may not be applicable to small-sized arrays. This paper proposes a deep learning assisted sound localization method using a small-sized microphone array constructed by two orthogonal first-order differential microphone arrays. An improved feature extraction scheme based on sound intensity estimation is also proposed by decoupling the correlation between sound pressure and particle velocity components in the whitening weighting construction to enhance the robustness of the time-frequency bin-wise sound intensity features. Simulation and real-world experimental results show that the proposed deep learning assisted approach can achieve higher spatial resolution and is superior to its state-of-the-art counterparts using the GCC-PHAT or sound intensity features for small-sized arrays in noisy and reverberant environments.