Stacked Temporal Attention: Improving First-person Action Recognition by Emphasizing Discriminative Clips

Stacked Temporal Attention: Improving First-person Action Recognition by Emphasizing Discriminative Clips
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
--
影响因子:
--
通讯作者:
Lijin Yang;Yifei Huang;Yusuke Sugano;Yoichi Sato
Lijin Yang;Yifei Huang;Yusuke Sugano;Yoichi Sato
中科院分区:
其他
文献类型:
--
作者:
Lijin Yang;Yifei Huang;Yusuke Sugano;Yoichi Sato

文献摘要

相似文献

第一人称动作识别是视频理解中的一项具有挑战性的任务。由于强烈的自我运动和有限的视野,第一人称视频中的许多背景或噪声帧可以在其学习过程中分散动作识别模型的注意力。为了编码更具区分性的特征,模型需要能够专注于视频中最相关的部分以进行动作识别。以前的工作探索,以解决这个问题,通过应用时间注意,但没有考虑到整个视频的全局上下文,这是确定相对重要的部分是至关重要的。在这项工作中,我们提出了一个简单而有效的堆叠时间注意力模块(STAM)来计算时间注意力的基础上,强调最具歧视性的功能的全球知识的剪辑。我们通过堆叠多个自我注意层来实现这一点。而不是天真的堆叠,这是实验证明是无效的,我们仔细设计的输入到每个自我注意力层,以便在生成时间注意力权重的过程中考虑视频的局部和全局上下文。实验表明,我们提出的STAM可以建立在大多数现有的骨干之上,并提高各种数据集的性能。
First-person action recognition is a challenging task in video understanding. Because of strong ego-motion and a limited field of view, many backgrounds or noisy frames in a first-person video can distract an action recognition model during its learning process. To encode more discriminative features, the model needs to have the ability to focus on the most relevant part of the video for action recognition. Previous works explored to address this problem by applying temporal attention but failed to consider the global context of the full video, which is critical for determining the relatively significant parts. In this work, we propose a simple yet effective Stacked Temporal Attention Module (STAM) to compute temporal attention based on the global knowledge across clips for emphasizing the most discriminative features. We achieve this by stacking multiple self-attention layers. Instead of naive stacking, which is experimentally proven to be ineffective, we carefully design the input to each self-attention layer so that both the local and global context of the video is considered during generating the temporal attention weights. Experiments demonstrate that our proposed STAM can be built on top of most existing backbones and boost the performance in various datasets.