Annotation-Based Multimedia Summarization and Translation

Annotation-Based Multimedia Summarization and Translation
复制标题

基于注释的多媒体摘要与翻译

DOI:
10.3115/1072228.1072326
复制
发表时间:
2002
期刊:
--
影响因子:
--
通讯作者:
Mitsuhiro Yoneoka
Mitsuhiro Yoneoka
中科院分区:
--
文献类型:
--
作者:
K. Nagao;S. Ohira;Mitsuhiro Yoneoka

文献摘要

被引文献

相似文献

本文介绍了多媒体注释技术及其在视频摘要和翻译中的应用。我们的注释工具允许用户轻松创建注释,包括语音转录,视频场景描述和视觉/听觉对象描述。语音转录模块能够进行多语种口语识别。视频场景描述由视频片段中每个场景的半自动检测的关键帧和场景的时间代码组成。通过对视频场景中的人和物体进行跟踪和交互命名来创建视觉对象描述。多媒体注释中的文本数据使用语言注释在句法和语义上结构化。所提出的多媒体摘要工作在多模态文档上,该文档由视频、场景的关键帧和场景的抄本组成。多媒体翻译自动生成不同语言的多媒体内容的多个版本。
This paper presents techniques for multimedia annotation and their application to video summarization and translation. Our tool for annotation allows users to easily create annotation including voice transcripts, video scene descriptions, and visual/auditory object descriptions. The module for voice transcription is capable of multilingual spoken language identification and recognition. A video scene description consists of semi-automatically detected keyframes of each scene in a video clip and time codes of scenes. A visual object description is created by tracking and interactive naming of people and objects in video scenes. The text data in the multimedia annotation are syntactically and semantically structured using linguistic annotation. The proposed multimedia summarization works upon a multimodal document that consists of a video, keyframes of scenes, and transcripts of the scenes. The multimedia translation automatically generates several versions of multimedia content in different languages.