Two-Stream 3-D convNet Fusion for Action Recognition in Videos With Arbitrary Size and Length

Two-Stream 3-D convNet Fusion for Action Recognition in Videos With Arbitrary Size and Length
复制标题

用于任意大小和长度视频中的动作识别的双流 3-D 卷积网络融合

DOI:
10.1109/tmm.2017.2749159
复制
发表时间:
2018-03-01
影响因子:
7.3
通讯作者:
Liu, Xianglong
Liu, Xianglong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wang, Xuanhan;Gao, Lianli;Liu, Xianglong

文献摘要

被引文献

相似文献

3-D卷积神经网络(3-D-convNets)最近被提出用于视频中的动作识别,并取得了可喜的成果。然而,现有的3-D-convNets具有两个可能降低视频分析质量的“人为”要求:1)它需要固定大小的(例如,112 × 112)输入视频;以及2)大多数3-D-convNet需要固定长度的输入(即,具有固定帧数的视频镜头)。为了解决这些问题,我们提出了一个名为Two-stream 3-D-convNet Fusion的端到端流水线,它可以使用多个特征识别任意大小和长度的视频中的人类动作。具体来说,我们将视频分解为空间和时间镜头。通过将镜头序列作为输入,每个流都使用具有长短期记忆(LSTM)或CNN-E模型的时空金字塔池化(STPP)convNet来实现,其中softmax分数通过后期融合来组合。我们设计了STPP convNet来为每个可变大小的镜头提取等维描述,并采用LSTM/CNN-E模型使用这些时变描述来学习输入视频的全局描述。有了这些优点,我们的方法应该改进所有基于3-D CNN的视频分析方法。我们经验性地评估了我们的视频动作识别方法,实验结果表明,我们的方法在三个标准基准数据集(UCF 101,HMDB 51和ACT数据集)上的性能优于最先进的方法(基于2-D和3-D)。
3-D convolutional neural networks (3-D-convNets) have been very recently proposed for action recognition in videos, and promising results are achieved. However, existing 3-D-convNets has two "artificial" requirements that may reduce the quality of video analysis: 1) It requires a fixed-sized (e.g., 112 x 112) input video; and 2) most of the 3-D-convNets require a fixed-length input (i.e., video shots with fixed number of frames). To tackle these issues, we propose an end-to-end pipeline named Two-stream 3-D-convNet Fusion, which can recognize human actions in videos of arbitrary size and length using multiple features. Specifically, we decompose a video into spatial and temporal shots. By taking a sequence of shots as input, each stream is implemented using a spatial temporal pyramid pooling (STPP) convNet with a long short-term memory (LSTM) or CNN-E model, softmax scores of which are combined by a late fusion. We devise the STPP convNet to extract equal-dimensional descriptions for each variable-size shot, and we adopt the LSTM/CNN-E model to learn a global description for the input video using these time-varying descriptions. With these advantages, our method should improve all 3-D CNN-based video analysis methods. We empirically evaluate our method for action recognition in videos and the experimental results show that our method outperforms the state-of-the-art methods (both 2-D and 3-D based) on three standard benchmark datasets (UCF101, HMDB51 and ACT datasets).