TubeR: Tubelet Transformer for Video Action Detection

TubeR: Tubelet Transformer for Video Action Detection
复制标题

DOI:
10.1109/cvpr52688.2022.01323
复制
发表时间:
2021-04
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Jiaojiao Zhao;Yanyi Zhang;Xinyu Li;Hao Chen;Shuai Bing;Mingze Xu;Chunhui Liu;Kaustav Kundu;Yuanjun Xiong;Davide Modolo;I. Marsic;Cees G. M. Snoek;Joseph Tighe
Jiaojiao Zhao;Yanyi Zhang;Xinyu Li;Hao Chen;Shuai Bing;Mingze Xu;Chunhui Liu;Kaustav Kundu;Yuanjun Xiong;Davide Modolo;I. Marsic;Cees G. M. Snoek;Joseph Tighe
中科院分区:
其他
文献类型:
--
作者:
Jiaojiao Zhao;Yanyi Zhang;Xinyu Li;Hao Chen;Shuai Bing;Mingze Xu;Chunhui Liu;Kaustav Kundu;Yuanjun Xiong;Davide Modolo;I. Marsic;Cees G. M. Snoek;Joseph Tighe

文献摘要

相似文献

我们提出TubeR:一个时空视频动作检测的简单解决方案。与现有依赖于离线演员检测器或手工设计的演员位置假设(如提案或锚)的方法不同,我们提出通过同时执行动作定位和从单个表示进行识别来直接检测视频中的动作小管。TubeR学习了一组管状查询,并利用管状注意模块对视频片段的动态时空特性进行建模,与在时空空间中使用角色位置假设相比,这有效地增强了模型的能力。对于包含过渡状态或场景变化的视频,我们提出了上下文感知分类头,利用短期和长期上下文加强动作分类,并提出了动作切换回归头,以检测精确的时间动作程度。TubeR直接生产可变长度的动作tubelet,甚至可以为长视频剪辑保持良好的效果。TubeR在常用的动作检测数据集AVA、UCF101-24和JHMDB51-21上优于以前的最先进技术。代码将在GluonCV(https://cv.gluon.ai/)上提供。
We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an offline actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in a video by simultaneously performing action localization and recognition from a single representation. TubeR learns a set of tubelet-queries and utilizes a tubelet-attention module to model the dynamic spatio-temporal nature of a video clip, which effectively reinforces the model capacity compared to using actor-positional hypotheses in the spatio-temporal space. For videos containing transitional states or scene changes, we propose a context aware classification head to utilize short-term and long-term context to strengthen action classification, and an action switch regression head for detecting the precise temporal action extent. TubeR directly produces action tubelets with variable lengths and even maintains good results for long video clips. TubeR outperforms the previous state-of-the-art on commonly used action detection datasets AVA, UCF101-24 and JHMDB51-21. Code will be available on GluonCV(https://cv.gluon.ai/).