AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation

AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation
复制标题

DOI:
10.1007/s11263-022-01702-9
复制
发表时间:
2022-10-28
影响因子:
19.5
通讯作者:
Ngan Le
Ngan Le
中科院分区:
计算机科学2区
文献类型:
--
作者:
Khoa Vo;Truong, Sang;Ngan Le

文献摘要

被引文献

相似文献

时间动作建议生成是一项具有挑战性的任务,它需要在未经裁剪的视频中定位动作间隔。直觉,我们作为人类,通过演员,相关对象和周围环境之间的相互作用来感知动作。尽管TAPG取得了重大进展,但绝大多数现有方法通过将骨干网络作为黑盒应用于给定视频而忽略了人类感知过程的上述原理。在本文中,我们提出了一个多模态表示网络,即演员对象环境交互网络(AOE-Net),这些相互作用的模型。我们的AOE-Net由两个模块组成,即,基于感知的多模态表示(PMR)和边界匹配模块(BMM)。此外,我们在PMR中引入自适应注意机制(AAM),只关注主要参与者(或相关对象),并对它们之间的关系进行建模。PMR模块通过视觉语言特征表示每个视频片段,其中主要演员和周围环境由视觉信息表示,而相关对象通过图像-文本模型由语言特征描述。BMM模块处理作为其输入的视觉语言特征序列,并生成行动建议。在ActivityNet-1.3和THUMOS-14数据集上进行的全面实验和广泛的消融研究表明,我们提出的AOE-Net优于以前最先进的方法,对于TAPG和时间动作检测都具有显着的性能和泛化能力。为了证明AOE-Net的鲁棒性和有效性,我们进一步对以自我为中心的视频(即EPIC-KITCHENS 100数据集)进行了消融研究。我们的源代码是公开的。
Temporal action proposal generation (TAPG) is a challenging task, which requires localizing action intervals in an untrimmed video. Intuitively, we as humans, perceive an action through the interactions between actors, relevant objects, and the surrounding environment. Despite the significant progress of TAPG, a vast majority of existing methods ignore the aforementioned principle of the human perceiving process by applying a backbone network into a given video as a black-box. In this paper, we propose to model these interactions with a multi-modal representation network, namely, Actors-Objects-Environment Interaction Network (AOE-Net). Our AOE-Net consists of two modules, i.e., perception-based multi-modal representation (PMR) and boundary-matching module (BMM). Additionally, we introduce adaptive attention mechanism (AAM) in PMR to focus only on main actors (or relevant objects) and model the relationships among them. PMR module represents each video snippet by a visual-linguistic feature, in which main actors and surrounding environment are represented by visual information, whereas relevant objects are depicted by linguistic features through an image-text model. BMM module processes the sequence of visual-linguistic features as its input and generates action proposals. Comprehensive experiments and extensive ablation studies on ActivityNet-1.3 and THUMOS-14 datasets show that our proposed AOE-Net outperforms previous state-of-the-art methods with remarkable performance and generalization for both TAPG and temporal action detection. To prove the robustness and effectiveness of AOE-Net, we further conduct an ablation study on egocentric videos, i.e. EPIC-KITCHENS 100 dataset. Our source code is publicly available at .