Weakly supervised learning of actions from transcripts

Weakly supervised learning of actions from transcripts
复制标题

DOI:
10.1016/j.cviu.2017.06.004
复制
发表时间:
2017-10-01
影响因子:
4.5
通讯作者:
Gall, Juergen
Gall, Juergen
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kuehne, Hilde;Richard, Alexander;Gall, Juergen

文献摘要

被引文献

相似文献

我们提出了一种弱监督学习的人的行动,从视频传输的方法。我们的系统是基于这样的想法,即给定一个输入数据序列和一个成绩单,即动作在视频中发生的顺序列表,可以推断视频流中的动作,并学习相关的动作模型,而不需要任何基于帧的注释。从手头的转录信息开始,我们根据预期操作的数量均匀地分割给定的数据序列。然后,我们通过最大化训练视频序列由动作模型生成的概率来学习每个类的动作模型,给定由成绩单定义的序列顺序。学习的模型可以用于在时间上分割具有或不具有转录本的不可见视频。此外,推断的片段可以用作训练高级完全监督模型的起点。我们在四个不同的活动数据集上评估了我们的方法,即Hollywood Extended,MPII Cooking,Breakfast和CRIM13。它表明,所提出的系统能够将脚本动作与视频数据对齐,学习模型对数据集中的动作进行定位和分类,并且它们优于任何当前最先进的将转录本与视频数据对齐的方法。(C)2017爱思唯尔公司All rights reserved.
We present an approach for weakly supervised learning of human actions from video transcriptions. Our system is based on the idea that, given a sequence of input data and a transcript, i.e. a list of the order the actions occur in the video, it is possible to infer the actions within the video stream and to learn the related action models without the need for any frame-based annotation. Starting from the transcript information at hand, we split the given data sequences uniformly based on the number of expected actions. We then learn action models for each class by maximizing the probability that the training video sequences are generated by the action models given the sequence order as defined by the transcripts. The learned model can be used to temporally segment an unseen video with or without transcript. Additionally, the inferred segments can be used as a starting point to train high-level fully supervised models.We evaluate our approach on four distinct activity datasets, namely Hollywood Extended, MPII Cooking, Breakfast and CRIM13. It shows that the proposed system is able to align the scripted actions with the video data, that the learned models localize and classify actions in the datasets, and that they outperform any current state-of-the-art approach for aligning transcripts with video data. (C) 2017 Elsevier Inc. All rights reserved.