GLNet: Global Local Network for Weakly Supervised Action Localization

GLNet: Global Local Network for Weakly Supervised Action Localization
复制标题

GLNet:弱监督动作本地化的全球局部网络

DOI:
10.1109/tmm.2019.2959425
复制
发表时间:
2020-10
影响因子:
7.3
通讯作者:
Nong Sang
Nong Sang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Shiwei Zhang;Lin Song;Changxin Gao;Nong Sang

文献摘要

参考文献

相似文献

在本文中,我们解决了弱监督时空动作定位的挑战性问题,对于该问题,在训练期间只有视频级别的动作标签可用。为了解决这一问题,我们提出了一种端到端的全球局域网络(GLNet)来同时在空间和时间空间上预测概率分布。提出的GLNet模型包括两个关键组件:局部空间模块和全局时间模块。局部空间模块旨在通过编码短期时间信息来预测帧级别的空间分布。特别是,我们提出了一种区域行为网络(RAN)来从预先计算的穷举提案中选择目标区域框。全局时间模块可以通过长期的时间结构建模来预测时间分布。具体地说,我们在几个片段的基础上设计了一种时间融合和激励结构,并通过稀疏损失函数进行训练。因此,所提出的GLNet模型能够以端到端的方式执行时空动作定位。我们在J-HMDB和UCF101-24数据集上评估了GLNet的性能。实验结果表明,GLNet在帧平均精度(MAP)和视频图(分别称为帧图和视频图)方面明显优于其他弱监督方法,甚至是一些完全监督方法。
In this paper, we address the challenging problem of weakly supervised spatio-temporal action localization for which only video-level action labels are available during training. To solve this problem, we propose an end-to-end Global Local Network (GLNet) to predict the probability distribution simultaneously in both spatial and temporal space. The proposed GLNet model includes two key components: a local spatial module and a global temporal module. The local spatial module aims to predict the frame-level spatial distribution by encoding short-term temporal information. In particular, we propose a Region Actionness Network (RAN) to select the target region boxes from the precomputed exhaustive proposals. The global temporal module can predict temporal distribution by a long-term temporal structure modelling. Specifically, we design a temporal fusion-and-excitation architecture on the top of several clips, and trained by a sparse loss function. Therefore, the proposed GLNet model can perform spatio-temporal action localization in an end-to-end manner. We evaluate the performance of GLNet on the J-HMDB and UCF101-24 datasets. The experimental results demonstrate GLNet achieves a significant margin against other state-of-the-art weakly supervised methods and even some fully supervised methods in terms of frame mean Average Precision (mAP) and the video mAP (called frame-mAP and video-mAP, respectively).
DOI: 10.1109/iccv.2017.473
发表时间: 2017-04
期刊: 2017 IEEE International Conference on Computer Vision (ICCV)
影响因子: --
作者:
Suman Saha;Gurkirt Singh;Fabio Cuzzolin
通讯作者: Suman Saha;Gurkirt Singh;Fabio Cuzzolin
DOI: 10.1109/iccv.2017.82
发表时间: 2017-10
期刊: 2017 IEEE International Conference on Computer Vision (ICCV)
影响因子: --
作者:
K. Soomro;M. Shah
通讯作者: K. Soomro;M. Shah
DOI: 10.1109/iccv.2017.476
发表时间: 2017-07
期刊: 2017 IEEE International Conference on Computer Vision (ICCV)
影响因子: --
作者:
P. Mettes;Cees G. M. Snoek
通讯作者: P. Mettes;Cees G. M. Snoek
DOI: 10.1109/tpami.2012.28
发表时间: 2012-11-01
影响因子: 23.6
作者:
Alexe, Bogdan;Deselaers, Thomas;Ferrari, Vittorio
通讯作者: Ferrari, Vittorio
DOI: 10.1109/tmm.2018.2814346
发表时间: 2017-11
影响因子: 7.3
作者:
Thanh-Toan Do;Tuan Hoang;Victor Pomponiu;Yiren Zhou;Zhao Chen;Ngai-Man Cheung;D. Koh;Aaron Tan
通讯作者: Thanh-Toan Do;Tuan Hoang;Victor Pomponiu;Yiren Zhou;Zhao Chen;Ngai-Man Cheung;D. Koh;Aaron Tan