Multi-modal speaker diarization of real-world meetings using compressed-domain video features

Multi-modal speaker diarization of real-world meetings using compressed-domain video features
复制标题

使用压缩域视频功能对现实世界会议进行多模式发言人分类

DOI:
10.1109/icassp.2009.4960522
复制
发表时间:
2009
期刊:
2009 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
Chuohao Yeo
Chuohao Yeo
中科院分区:
--
文献类型:
--
作者:
G. Friedland;H. Hung;Chuohao Yeo

文献摘要

被引文献

相似文献

发言者日记最初被定义为在没有任何其他先验知识的情况下确定“谁在什么时候发言”的任务。下面的文章展示了一种多模态方法,我们通过将标准声学特征(MFCC)与压缩域视频特征相结合来改进最先进的扬声器日志化系统。该方法在超过4.5小时的公开AMI会议数据集上进行了评估,其中包含了人们站起来和走出房间等挑战。与最先进的仅音频基线相比,我们显示出相对于扬声器错误率(21%DER)约34%的一致改善。
Speaker diarization is originally defined as the task of determining “who spoke when” given an audio track and no other prior knowledge of any kind. The following article shows a multi-modal approach where we improve a state-of-the-art speaker diarization system by combining standard acoustic features (MFCCs) with compressed domain video features. The approach is evaluated on over 4.5 hours of the publicly available AMI meetings dataset which contains challenges such as people standing up and walking out of the room. We show a consistent improvement of about 34% relative in speaker error rate (21% DER) compared to a state-of-the-art audio-only baseline.