Multi-modal dialog scene detection using hidden Markov models for content-based multimedia indexing

Multi-modal dialog scene detection using hidden Markov models for content-based multimedia indexing
复制标题

DOI:
10.1023/a:1011395131992
复制
发表时间:
2001-06-01
影响因子:
3.6
通讯作者:
Wolf, W
Wolf, W
中科院分区:
计算机科学4区
文献类型:
--
作者:
Alatan, AA;Akansu, AN;Wolf, W

文献摘要

被引文献

相似文献

使用新颖的基于隐马尔可夫模型(HMM)的方法将一类视听数据(小说娱乐:电影、电视剧)分割成包含对话的场景。每个镜头都使用音轨(通过语音、静音和音乐分类)和视觉内容(面部和位置信息)进行分类。这种基于镜头的分类的结果是一个视听标记,HMM 状态图将使用该标记来实现场景分析。在使用圆形和从左到右的 HMM 拓扑进行仿真后,我们发现两者在多模态输入下都表现得非常好。此外,对于圆形拓扑,不同训练集和观察集之间的比较表明,音频和面部信息一起给出了不同观察集之间最一致的结果。
A class of audio-visual data (fiction entertainment: movies, TV series) is segmented into scenes, which contain dialogs, using a novel hidden Markov model-based (HMM) method. Each shot is classified using both audio track (via classification of speech, silence and music) and visual content (face and location information). The result of this shot-based classification is an audio-visual token to be used by the HMM state diagram to achieve scene analysis. After simulations with circular and left-to-right HMM topologies, it is observed that both are performing very good with multi-modal inputs. Moreover, for circular topology, the comparisons between different training and observation sets show that audio and face information together gives the most consistent results among different observation sets.