A SCALABLE FRAMEWORK FOR MULTIPLE SPEAKER LOCALIZATION AND TRACKING

A SCALABLE FRAMEWORK FOR MULTIPLE SPEAKER LOCALIZATION AND TRACKING
复制标题

用于多说话者定位和跟踪的可扩展框架

DOI:
--
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
N. Madhu
N. Madhu
中科院分区:
--
文献类型:
--
作者:
N. Madhu

文献摘要

被引文献

相似文献

在本文中,我们提出了一种新的,可扩展的方法来定位和跟踪多个扬声器使用麦克风阵列。该方法能够定位源在非竞争和并发的情况下,是基于在短时间离散频域(STFD)的语音的不相交性。该算法对STFD中的窄带定位成本函数进行操作,并通过将高斯混合(MoG)拟合应用于窄带定位估计来产生每个时间帧B的说话者活动的估计。所提出的方法的优点是多方面的:它允许我们使用一个粗糙的搜索网格的成本函数评估,而不影响的位置精度;它允许一个软决策的数量和位置的源在飞行中;该框架是可扩展的多阵列系统;它也可以作为一个基础框架的增强算法。原则上,这种方法并不特定于扬声器,只要它们表现出一定的时间和频谱不相交性,它就适用于任何源组合。
In this paper we present a novel, scalable approach to the localization and tracking of multiple speakers using microphone arrays. The approach is capable of localizing sources both in non-competing and in concurrent situations, and is based on the disjointness of speech in the short-time discrete frequency domain (STFD). The algorithm operates on a narrowband localization cost function in the STFD and yields an estimate of the speaker activity per time frame b by applying a Mixture of Gaussians (MoG) fit to the narrowband localization estimates. The advantages of the proposed method are manifold: it allows us to use a coarser search grid for the cost function evaluation, without compromising on the location accuracy; it allows for a soft-decision on the number and position of sources on the fly; the framework is scalable to multi-array systems; and it can also serve as a base framework for enhancement algorithms. In principle, this approach is not specific to speakers and will work for any source combination provided they exhibit some temporal and spectral disjointness.