Spatiotemporal Residual Networks for Video Action Recognition

Spatiotemporal Residual Networks for Video Action Recognition
复制标题

DOI:
--
复制
发表时间:
2016-11
期刊:
--
影响因子:
--
通讯作者:
Christoph Feichtenhofer;A. Pinz;Richard P. Wildes
Christoph Feichtenhofer;A. Pinz;Richard P. Wildes
中科院分区:
其他
文献类型:
--
作者:
Christoph Feichtenhofer;A. Pinz;Richard P. Wildes

文献摘要

被引文献

相似文献

双流卷积网络(ConvNets)在视频中的人体动作识别方面表现出了很强的性能。最近,剩余网络(ResNets)作为一种新的技术出现,用于训练极深层次的体系结构。在本文中,我们引入了时空ResNet作为这两种方法的结合。我们的新体系结构通过两种方式引入剩余连接,从而将ResNet推广到时空领域。首先,我们在双流架构的外观和运动路径之间注入残余连接,以允许两个流之间的时空交互。其次,我们将预先训练好的图像ConvNet转换为时空网络,给这些网络配备可学习的卷积过滤器,这些过滤器被初始化为时间残差连接,并在时间上对相邻的特征映射进行操作。这种方法随着模型深度的增加而缓慢地增加时空感受野,并自然地融合了图像ConvNet的设计原则。整个模型被端到端地训练,以允许对复杂的时空特征进行分层学习。我们使用两个广泛使用的动作识别基准对我们的新时空ResNet进行了评估,其中它超过了以前的最先进水平。
Two-stream Convolutional Networks (ConvNets) have shown strong performance for human action recognition in videos. Recently, Residual Networks (ResNets) have arisen as a new technique to train extremely deep architectures. In this paper, we introduce spatiotemporal ResNets as a combination of these two approaches. Our novel architecture generalizes ResNets for the spatiotemporal domain by introducing residual connections in two ways. First, we inject residual connections between the appearance and motion pathways of a two-stream architecture to allow spatiotemporal interaction between the two streams. Second, we transform pretrained image ConvNets into spatiotemporal networks by equipping these with learnable convolutional filters that are initialized as temporal residual connections and operate on adjacent feature maps in time. This approach slowly increases the spatiotemporal receptive field as the depth of the model increases and naturally integrates image ConvNet design principles. The whole model is trained end-to-end to allow hierarchical learning of complex spatiotemporal features. We evaluate our novel spatiotemporal ResNet using two widely used action recognition benchmarks where it exceeds the previous state-of-the-art.