Spatio-Temporal Vector of Locally Max Pooled Features for Action Recognition in Videos

Spatio-Temporal Vector of Locally Max Pooled Features for Action Recognition in Videos
复制标题

DOI:
10.1109/cvpr.2017.341
复制
发表时间:
2017-07
期刊:
2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Ionut Cosmin Duta;B. Ionescu;K. Aizawa;N. Sebe
Ionut Cosmin Duta;B. Ionescu;K. Aizawa;N. Sebe
中科院分区:
其他
文献类型:
--
作者:
Ionut Cosmin Duta;B. Ionescu;K. Aizawa;N. Sebe

文献摘要

被引文献

相似文献

我们介绍了局部最大池化特征的时空向量(ST-VLMPF),这是一种专门为局部深度特征编码而设计的基于超级向量的编码方法。所提出的方法解决了视频理解的一个重要问题:如何构建一个在整个视频中包含CNN特征的视频表示。特征分配是在两个层次上进行的,通过使用相似性和时空信息。对于每个任务,我们都会构建一个特定的编码,重点关注深度特征的本质,目标是从网络的最高神经元激活中捕获最高的特征响应。我们的ST-VLMPF显然提供了比一些最广泛使用和功能强大的编码方法(改进的Fisher矢量和局部聚合描述符矢量)更可靠的视频表示,同时保持了较低的计算复杂度。我们在三个动作识别数据集上进行了实验:HMDB 51,UCF 50和UCF 101。我们的管道获得了最先进的结果。
We introduce Spatio-Temporal Vector of Locally Max Pooled Features (ST-VLMPF), a super vector-based encoding method specifically designed for local deep features encoding. The proposed method addresses an important problem of video understanding: how to build a video representation that incorporates the CNN features over the entire video. Feature assignment is carried out at two levels, by using the similarity and spatio-temporal information. For each assignment we build a specific encoding, focused on the nature of deep features, with the goal to capture the highest feature responses from the highest neuron activation of the network. Our ST-VLMPF clearly provides a more reliable video representation than some of the most widely used and powerful encoding approaches (Improved Fisher Vectors and Vector of Locally Aggregated Descriptors), while maintaining a low computational complexity. We conduct experiments on three action recognition datasets: HMDB51, UCF50 and UCF101. Our pipeline obtains state-of-the-art results.