Associating video frames with text

Associating video frames with text
复制标题

将视频帧与文本关联

DOI:
--
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
H. Wactlar
H. Wactlar
中科院分区:
--
文献类型:
--
作者:
Pinar Duygulu;H. Wactlar

文献摘要

被引文献

相似文献

在本研究中,提出了视觉和文本数据的集成来解决视频帧和相关文本之间的对应问题,以便用更可靠的标签和描述来注释视频帧。从视频帧中提取的视觉特征链接到使用联合统计从音频记录中获得的文本。结果表明,使用这种方法可以为视频帧提供更好的注释,随后可用于提高基于文本的查询的性能。所提出的方法将被整合到卡内基梅隆大学的 Infomedia 数字视频图书馆项目中。
In this study, integration of visual and textual data is proposed to solve the correspondence problem between video frames and associated text in order to annotate video frames with more reliable labels and descriptions. Visual features extracted from video frames are linked to text that is obtained from the audio transcripts using joint statistics. The results show that using this approach it is possible to have better annotations for video frames that can be later used to improve the performance of text based queries. The proposed approach will be integrated into the Informedia Digital Video Library Project at Carnegie Mellon University.