Delving Deeper into Convolutional Networks for Learning Video Representations

Delving Deeper into Convolutional Networks for Learning Video Representations
复制标题

DOI:
--
复制
发表时间:
2015-11
期刊:
CoRR
影响因子:
--
通讯作者:
Nicolas Ballas;L. Yao;C. Pal;Aaron C. Courville
Nicolas Ballas;L. Yao;C. Pal;Aaron C. Courville
中科院分区:
其他
文献类型:
--
作者:
Nicolas Ballas;L. Yao;C. Pal;Aaron C. Courville

文献摘要

被引文献

相似文献

我们提出了一种使用门控循环单元循环网络 (GRU) 从我们称为“感知”的中间视觉表示中学习视频中时空特征的方法。我们的方法依赖于从在大型 ImageNet 数据集上训练的深度卷积网络的各个级别中提取的感知。虽然高级感知包含高度辨别性的信息,但它们往往具有较低的空间分辨率。另一方面,低级感知保留了更高的空间分辨率,我们可以从中建模更精细的运动模式。使用低级感知可以产生高维视频表示。为了减轻这种影响并控制模型参数数量,我们引入了 GRU 模型的一种变体,它利用卷积运算来强制模型单元的稀疏连接并在输入空间位置之间共享参数。我们根据经验验证了我们在人类动作识别和视频字幕任务上的方法。特别是,我们使用更简单的文本解码器模型且无需额外的 3D CNN 功能,在 YouTube2Text 数据集上获得了相当于最新技术的结果。
We propose an approach to learn spatio-temporal features in videos from intermediate visual representations we call "percepts" using Gated-Recurrent-Unit Recurrent Networks (GRUs).Our method relies on percepts that are extracted from all level of a deep convolutional network trained on the large ImageNet dataset. While high-level percepts contain highly discriminative information, they tend to have a low-spatial resolution. Low-level percepts, on the other hand, preserve a higher spatial resolution from which we can model finer motion patterns. Using low-level percepts can leads to high-dimensionality video representations. To mitigate this effect and control the model number of parameters, we introduce a variant of the GRU model that leverages the convolution operations to enforce sparse connectivity of the model units and share parameters across the input spatial locations. We empirically validate our approach on both Human Action Recognition and Video Captioning tasks. In particular, we achieve results equivalent to state-of-art on the YouTube2Text dataset using a simpler text-decoder model and without extra 3D CNN features.