A Mixed-State I-Particle Filter for Multi-Camera Speaker Tracking

A Mixed-State I-Particle Filter for Multi-Camera Speaker Tracking
复制标题

用于多摄像机扬声器跟踪的混合状态 I 粒子滤波器

DOI:
--
复制
发表时间:
2003
期刊:
IEEE International Conference on Computer Vision
影响因子:
--
通讯作者:
J. Odobez
J. Odobez
中科院分区:
--
文献类型:
--
作者:
D. Gática;Guillaume Lathoud;I. McCowan;J. Odobez

文献摘要

被引文献

相似文献

跟踪多方对话中的发言者是实现会议自动分析的重要一步。本文提出了一种在多传感器会议室中跟踪视听说话人的概率方法。该算法通过混合状态重要性粒子过滤器融合来自三个未校准摄像头和麦克风阵列的信息,允许整合AV流以利用每种模式的互补特征。我们的方法依赖于几个原则。首先,使用混合状态空间公式定义摄像机切换的产生式模型。其次,使用AV定位信息来定义重要性采样函数,该采样函数引导粒子过滤器的搜索过程朝向可能包含真实配置(扬声器)的配置空间的区域。最后,测量过程集成了形状、颜色和音频观察。我们表明,不完美模态的原则组合导致了一种算法,该算法自动初始化和跟踪参与真实对话的说话人,可靠地在摄像机和参与者之间切换。
Tracking speakers in multi-party conversations represents an important step towards automatic analysis of meetings. In this paper, we present a probabilistic method for audio-visual (AV) speaker tracking in a multi-sensor meeting room. The algorithm fuses information coming from three uncalibrated cameras and a microphone array via a mixed-state importance particle filter, allowing for the integration of AV streams to exploit the complementary features of each modality. Our method relies on several principles. First, a mixed state space formulation is used to define a generative model for camera switching. Second, AV localization information is used to define an importance sampling function, which guides the search process of a particle filter towards regions of the configuration space likely to contain the true configuration (a speaker). Finally, the measurement process integrates shape, color, and audio observations. We show that the principled combination of imperfect modalities results in an algorithm that automatically initializes and tracks speakers engaged in real conversations, reliably switching across cameras and between participants.