Detection of Inconsistency Between Subject and Speaker Based on the Co-occurrence of Lip Motion and Voice Towards Speech Scene Extraction from News Videos

Detection of Inconsistency Between Subject and Speaker Based on the Co-occurrence of Lip Motion and Voice Towards Speech Scene Extraction from News Videos
复制标题

DOI:
10.1109/ism.2011.56
复制
发表时间:
2011-12
期刊:
2011 IEEE International Symposium on Multimedia
影响因子:
--
通讯作者:
S. Kumagai;Keisuke Doman;Tomokazu Takahashi;Daisuke Deguchi;I. Ide;H. Murase
S. Kumagai;Keisuke Doman;Tomokazu Takahashi;Daisuke Deguchi;I. Ide;H. Murase
中科院分区:
其他
文献类型:
--
作者:
S. Kumagai;Keisuke Doman;Tomokazu Takahashi;Daisuke Deguchi;I. Ide;H. Murase

文献摘要

相似文献

我们提出了一种方法来检测主题和扬声器之间的不一致性提取语音场景从新闻视频。新闻视频中的演讲场景包含丰富的多媒体信息,作为存档材料很有价值。为了从新闻视频中提取语音场景,存在使用人脸区域的位置和大小的方法。然而,仅用这种方法很难提取它们,因为新闻视频包含说话者不是主体的非语音场景,例如叙述场景。为了解决这个问题,我们提出了一种方法来区分语音场景和叙述场景的基础上,一个主体的嘴唇运动和扬声器的声音之间的共现。该方法使用嘴唇形状和程度的嘴唇开放的视觉特征表示一个主体的嘴唇运动,并使用语音音量和音素作为音频特征表示说话人的声音。然后,该方法区分语音场景和叙述场景的基础上,这些功能的相关性。我们报告的实验结果,在实验室条件下捕获的视频,也对实际的广播新闻视频。结果表明了该方法的有效性和研究目标的可行性。
We propose a method to detect the inconsistency between a subject and the speaker for extracting speech scenes from news videos. Speech scenes in news videos contain a wealth of multimedia information, and are valuable as archived material. In order to extract speech scenes from news videos, there is an approach that uses the position and size of a face region. However, it is difficult to extract them with only such approach, since news videos contain non-speech scenes where the speaker is not the subject, such as narrated scenes. To solve this problem, we propose a method to discriminate between speech scenes and narrated scenes based on the co-occurrence between a subject's lip motion and the speaker's voice. The proposed method uses lip shape and degree of lip opening as visual features representing a subject's lip motion, and uses voice volume and phoneme as audio feature representing a speaker's voice. Then, the proposed method discriminates between speech scenes and narrated scenes based on the correlations of these features. We report the results of experiments on videos captured in a laboratory condition and also on actual broadcast news videos. Their results showed the effectiveness of our method and the feasibility of our research goal.