A fast cumulative Steered Response Power for multiple speaker detection and localization

A fast cumulative Steered Response Power for multiple speaker detection and localization
复制标题

用于多说话人检测和定位的快速累积转向响应功率

DOI:
--
复制
发表时间:
2013
期刊:
European Signal Processing Conference
影响因子:
--
通讯作者:
D. Klakow
D. Klakow
中科院分区:
--
文献类型:
--
作者:
Youssef Oualil;F. Faubel;D. Klakow

文献摘要

被引文献

相似文献

本文提出了一种新的方法,用于检测和定位多个扬声器使用麦克风阵列。在该框架中,经典的转向响应功率(SRP)技术与一种新颖的两步搜索策略相结合,以降低计算成本。这里采用的方法通过以下方式执行定位:1)使用每个广义互相关(GCC)函数提供的空间信息将搜索空间减少到几个可能包含源的子空间。从这些中,最可能的区域被提取为使累积SRP最大化的子空间。然后,2)在约简空间中使用经典搜索方法估计最佳源位置。使用无监督贝叶斯分类器进一步改进了噪声/说话人检测。在AV16.3语料库上的实验表明,该方法比经典的SRP方法快约47倍,而定位性能没有任何明显的下降。
This paper presents a novel approach for detecting and localizing multiple speakers using a microphone array. In this framework, the classical Steered Response Power (SRP) technique is combined with a novel two-step search strategy to reduce the computation cost. The approach taken here performs the localization by 1) using the spatial information provided by each Generalized Cross Correlation (GCC) function to reduce the search space to a few subspaces that are likely to contain a source. From these, the most likely region is extracted as the subspace that maximizes the Cumulative SRP. Then, 2) the optimal source location is estimated using the classical search approach in the reduced space. The noise/speaker detection is further improved using an unsupervised Bayesian classifier. Experiments on the AV16.3 corpus show that the proposed method is approximately 47 times faster than the classical SRP, without any noticeable degradation of the localization performance.