Ordered Pooling of Optical Flow Sequences for Action Recognition

Ordered Pooling of Optical Flow Sequences for Action Recognition
复制标题

DOI:
10.1109/wacv.2017.26
复制
发表时间:
2017-01
期刊:
2017 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Jue Wang;A. Cherian;F. Porikli
Jue Wang;A. Cherian;F. Porikli
中科院分区:
其他
文献类型:
--
作者:
Jue Wang;A. Cherian;F. Porikli

文献摘要

被引文献

相似文献

在长视频序列上训练卷积神经网络(CNN)在计算上是昂贵的,因为深度架构需要大量的内存和大量的参数。因此,视频帧的早期融合是一种标准技术,其中几个连续的帧首先聚集成一个紧凑的表示,然后作为输入样本馈送到CNN中。为此,最近提出了一种通过单个动态图像来表示一组连续RGB帧以捕获像素动态的摘要方法。在本文中,我们介绍了一种新的连续光流帧的有序表示作为一种替代,并认为这种表示捕获的动作动态比RGB帧更有效。我们提供了为什么这样的表示是更好的动作识别的直觉。我们在标准基准数据集上验证了我们的声明,并证明了使用流图像的摘要可以显著改善RGB帧,同时实现与UCF101和HMDB数据集上最先进的精度相当的精度。
Training of Convolutional Neural Networks (CNNs) on long video sequences is computationally expensive due to the substantial memory requirements and the massive number of parameters that deep architectures demand. Early fusion of video frames is thus a standard technique, in which several consecutive frames are first agglomerated into a compact representation, and then fed into the CNN as an input sample. For this purpose, a summarization approach that represents a set of consecutive RGB frames by a single dynamic image to capture pixel dynamics is proposed recently. In this paper, we introduce a novel ordered representation of consecutive optical flow frames as an alternative and argue that this representation captures the action dynamics more efficiently than RGB frames. We provide intuitions on why such a representation is better for action recognition. We validate our claims on standard benchmark datasets and demonstrate that using summaries of flow images lead to significant improvements over RGB frames while achieving accuracy comparable to the stateof-the-art on UCF101 and HMDB datasets.