Actor-Centered Representations for Action Localization in Streaming Videos

Actor-Centered Representations for Action Localization in Streaming Videos
复制标题

DOI:
10.1007/978-3-031-19839-7_5
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
Sathyanarayanan N. Aakur;Sudeep Sarkar
Sathyanarayanan N. Aakur;Sudeep Sarkar
中科院分区:
其他
文献类型:
--
作者:
Sathyanarayanan N. Aakur;Sudeep Sarkar

文献摘要

相似文献

识别和定位流媒体视频中的动作等事件感知任务对于扩展到现实世界的应用程序上下文至关重要。我们通过连续分层预测学习的概念来解决学习以演员为中心的再现的问题,以定位流视频中的动作,而不需要训练视频中对象的标签和轮廓。我们提出了一个框架驱动的概念,分层预测学习的constructator-centeredfeatures的注意力为基础的contextualization。其关键思想是,可预测的特征或对象不会引起注意,因此不会对感兴趣的动作做出贡献。在三个基准数据集上的实验表明,该方法只需一个训练时期,即,一次通过视频流。我们表明,所提出的方法优于无监督和弱监督基线,同时提供有竞争力的性能完全监督的方法。此外,我们将该模型扩展到多演员设置,以识别组活动,同时本地化的多个,似是而非的演员。我们还表明,它一般化到域外的数据有限的性能下降。
Event perception tasks such as recognizing and localizing actions in streaming videos are essential for scaling to real-world application contexts. We tackle the problem of learningactor-centeredrepresentations through the notion ofcontinual hierarchical predictive learningtolocalizeactions in streaming videoswithoutthe need for training labels and outlines for the objects in the video. We propose a framework driven by the notion of hierarchical predictive learning to constructactor-centeredfeatures by attention-based contextualization. The key idea is that predictable features or objects do not attract attention and hence do not contribute to the action of interest. Experiments on three benchmark datasets show that the approach can learn robust representations for localizing actionsusing only one epoch of training, i.e., a single pass through the streaming video. We show that the proposed approach outperforms unsupervised and weakly supervised baselines while offering competitive performance to fully supervised approaches. Additionally, we extend the model to multi-actor settings to recognize group activities while localizing the multiple, plausible actors. We also show that it generalizes to out-of-domain data with limited performance degradation.