Multimodal Speaker Segmentation in Presence of Overlapped Speech Segments

Multimodal Speaker Segmentation in Presence of Overlapped Speech Segments
复制标题

存在重叠语音段的多模态说话人分割

DOI:
--
复制
发表时间:
2008
期刊:
2008 Tenth IEEE International Symposium on Multimedia
影响因子:
--
通讯作者:
Shrikanth S. Narayanan
Shrikanth S. Narayanan
中科院分区:
--
文献类型:
--
作者:
Viktor Rozgic;Kyu Jeong Han;P. Georgiou;Shrikanth S. Narayanan

文献摘要

被引文献

相似文献

我们提出了一个多模态说话人分割算法,主要有两个贡献:第一,我们提出了一个隐马尔可夫模型架构,执行三个模态的融合:多摄像机系统的参与者定位,麦克风阵列的说话人定位,和说话人识别系统;第二,我们提出了一种新的方法,用于处理重叠的语音段,通过麦克风阵列观测的似然模型,使用多个在联合概率数据关联(JPDA)框架中的转向功率响应广义互相关相位变换(SPR-GCC-PHAT)函数的局部最大值。结果表明,该方法优于标准的说话人分割系统的基础上:(a)说话人识别和(B)麦克风阵列处理,数据集的重叠语音的显着部分(27.4%),分数高达94.4%的F-测量尺度。
We propose a multimodal speaker segmentation algorithm with two main contributions: First, we suggest a hidden Markov model architecture that performs fusion of the three modalities: a multi-camera system for participant localization, a microphone array for speaker localization, and a speaker identification system; second, we present a novel method for dealing with overlapped speech segments through a likelihood model of the microphone array observations that uses multiple local maxima of the Steered Power response generalized cross correlation phase transform (SPR-GCC-PHAT) function in the joint probabilistic data association (JPDA) framework. Results show that the proposed method outperforms standard speaker segmentation systems based on: (a) speaker identification and; (b) microphone array processing, for datasets with the significant portion (27.4%) of overlapped speech, and scores as high as 94.4% on the F-measure scale.