A SCALABLE FRAMEWORK FOR MULTIPLE SPEAKER LOCALIZATION AND TRACKING
A SCALABLE FRAMEWORK FOR MULTIPLE SPEAKER LOCALIZATION AND TRACKING
复制标题
用于多说话者定位和跟踪的可扩展框架
DOI:
--
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
N. Madhu
中科院分区:
文献类型:
--
作者:
N. Madhu
In this paper we present a novel, scalable approach to the localization and tracking of multiple speakers using microphone arrays. The approach is capable of localizing sources both in non-competing and in concurrent situations, and is based on the disjointness of speech in the short-time discrete frequency domain (STFD). The algorithm operates on a narrowband localization cost function in the STFD and yields an estimate of the speaker activity per time frame b by applying a Mixture of Gaussians (MoG) fit to the narrowband localization estimates. The advantages of the proposed method are manifold: it allows us to use a coarser search grid for the cost function evaluation, without compromising on the location accuracy; it allows for a soft-decision on the number and position of sources on the fly; the framework is scalable to multi-array systems; and it can also serve as a base framework for enhancement algorithms. In principle, this approach is not specific to speakers and will work for any source combination provided they exhibit some temporal and spectral disjointness.