课题基金 / 基金详情

Automatic Alignment of Textto-Video for Semantic Multimedia Analysis

Automatic Alignment of Textto-Video for Semantic Multimedia Analysis
用于语义多媒体分析的文本到视频的自动对齐
批准号:
252286362
负责人:
Professor Dr.-Ing. Rainer Stiefelhagen
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2014
资助国家:
德国
项目状态:
已结题
起止时间:
2013-12-31 至 2017-12-31

项目摘要

项目成果

Professor Dr.-Ing. Rainer Stiefelhagen的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
In this project, we aim to explore rich descriptions of video data (TV series and movies) which opens myriad possibilities for multimedia analysis, understanding and obtaining weak labels for popular computer vision tasks. We wish to focus on two forms of text -- plot synopses and books. The former, plots are obtained via crowdsourcing and describe the episode or movie in a summarized way. In contrast books (from which the video is adapted) provide detailed descriptions of the story and visual world the author wishes to portray.While text in the form of subtitles and transcripts has been successfully used to automate person identification [Everingham 2006] or obtain samples for action recognition [Laptev 2008], those text sources are limited in their potential for understanding or obtaining rich descriptions of the story.To use the plot synopses, we will first align the sentences of the synopsis to shots in the video (WP2). We propose to use anchors, primarily person-id to help guide the alignment. We aim to solve two main challenges associated with this task: possible non-linearity of the plot synopsis, and skipping of shots.In contrast to plot synopses, the first step we take in analyzing books is to align chapters and their corresponding video shots (WP3). We can expect that some dialogues in the books match the ones used in the video adaptation. This allows us to automatically identify characters and learn person models in a second step, and also facilitates fine-grained alignment within a chapter.The alignment can be improved by knowing more about the scene or objects present in the shots. We will investigate this interconnected behaviour of labels and anchors in WP4, first in an iterative manner, and then by jointly modeling the two tasks of obtaining weak labels and performing alignment.We divide the applications into two types: (i) obtaining labels from the text sources and (ii) video-related applications. From plot synopses, we will specifically aim to obtain weak labels for places or scenes (WP5-P1). We will also explore tasks such as Summarization, Indexing and Retrieval (WP5-P2). For example, a coherent video summary based on the story (rather than low-level features) can be generated by first running a text summarizer on the plot, followed by selection of the set of aligned to the retained sentences. Indexing the descriptions for keywords can also lead to easy browsing through the video. From books, we wish to exploit dialogs for obtaining supervision for person identification, and rich descriptions surrounding the dialogs to learn attributes for the characters, scenes and objects (WP5-P1). Another interesting application is to automatically find differences between books and their video adaptations (WP5-P2).
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1007/s13735-014-0065-9
发表时间: 2015-03
期刊: International Journal of Multimedia Information Retrieval
影响因子: 5.6
作者: [Makarand Tapaswi;M. Bäuml;R. Stiefelhagen]
通讯作者: Makarand Tapaswi;M. Bäuml;R. Stiefelhagen
ComPLetely Unsupervised Multimodal Character identification On TV series and movies
国内基金
海外基金
序列比对( Alignment)的随机分析与快速算法
  • 批准号:
    10271061
  • 项目类别:
    面上项目
  • 资助金额:
    16.5万元
  • 批准年份:
    2002
  • 负责人:
    沈世镒
  • 依托单位: