课题基金 / 基金详情

ComPLetely Unsupervised Multimodal Character identification On TV series and movies

ComPLetely Unsupervised Multimodal Character identification On TV series and movies
电视剧和电影中完全无监督的多模态角色识别
批准号:
316692988
负责人:
Professor Dr.-Ing. Rainer Stiefelhagen
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2016
资助国家:
德国
项目状态:
已结题
起止时间:
2015-12-31 至 2020-12-31

项目摘要

项目成果

Professor Dr.-Ing. Rainer Stiefelhagen的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Automatic character identification in multimedia videos is an extensive and challenging problem. Person identification serves as foundation and building block for many higher level video analysis tasks, for example semantic indexing, search and retrieval, interaction analysis and video summarization.The goal of this project is to exploit textual, audio and video information to automatically identify characters in TV series and movies without requiring any manual annotation for training character models. A fully automatic and unsupervised approach is especially appealing when considering the huge amount and growth of available multimedia data. Text, audio and video provide complementary cues to the identity of a person, and thus allow to better identify a person than from either modality alone.To this end, we will address three main research questions: unsupervised clustering of speech turns (i.e. speaker diarization) and face tracks in order to group similar tracks of the same person without prior labels or models; unsupervised identification by propagation of automatically generated weak labels from various sources of information (such as subtitles and speech transcripts); and multimodal fusion of acoustic, visual and textual cues at various levels of the identification pipeline.While there exist many generic approaches to unsupervised clustering, they are not adapted to heterogeneous audiovisual data (face tracks vs. speech turns) and do not perform as well on challenging TV series and movies content as they do on other controlled data. Our general approach is therefore to first over-cluster the data and make sure that clusters remain pure, before assigning names to these clusters in a second step. On top of domain specific improvements for either modality alone, we expect joint multimodal clustering to take advantage of three modalities and improve clustering performance over each modality alone.Then, unsupervised identification aims at assigning character names to clusters in a completely automatic manner (i.e. using only available information already present in the speech and video). In TV series and movies, character names are usually introduced and re-iterated throughout the video. We will detect and use addresser-addressee relationships in both speech transcripts (using named entity detection techniques) and video (using mouth movements, viewing direction and focus of attention of faces). This allows to assign names to some clusters, learn discriminative models and assign names to the remaining clusters.For evaluation, we will extend and further annotate a corpus of four TV series (57 episodes) and one movie series (8 movies), a total of about 50 hours of video. This diverse data covers different filming styles, type of stories and challenges contained in both video and audio. We will evaluate the different steps of this project on this corpus, and also make our annotations publicly available for other researchers in the field.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/cvpr.2018.00424
发表时间: 2017-10
期刊: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition
影响因子: --
作者: [Vivek Sharma;Ali Diba;D. Neven;M. S. Brown;L. Gool;R. Stiefelhagen]
通讯作者: Vivek Sharma;Ali Diba;D. Neven;M. S. Brown;L. Gool;R. Stiefelhagen
DOI: 10.1145/3343031.3351071
发表时间: 2019-10
期刊: Proceedings of the 27th ACM International Conference on Multimedia
影响因子: --
作者: [Veith Röthlingshöfer;Vivek Sharma;R. Stiefelhagen]
通讯作者: Veith Röthlingshöfer;Vivek Sharma;R. Stiefelhagen
DOI: 10.1109/fg47880.2020.00011
发表时间: 2020-04
期刊: 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020)
影响因子: --
作者: [Vivek Sharma;Makarand Tapaswi;M. Sarfraz;R. Stiefelhagen]
通讯作者: Vivek Sharma;Makarand Tapaswi;M. Sarfraz;R. Stiefelhagen
DOI: 10.1109/tbiom.2019.2947264
发表时间: 2020-04
期刊: IEEE Transactions on Biometrics, Behavior, and Identity Science
影响因子: --
作者: [Vivek Sharma;Makarand Tapaswi;M. Sarfraz;R. Stiefelhagen]
通讯作者: Vivek Sharma;Makarand Tapaswi;M. Sarfraz;R. Stiefelhagen
Automatic Alignment of Textto-Video for Semantic Multimedia Analysis
海外基金