3D Convolutional Neural Networks for Human Action Recognition

3D Convolutional Neural Networks for Human Action Recognition
复制标题

DOI:
10.1109/tpami.2012.59
复制
发表时间:
2013-01-01
影响因子:
23.6
通讯作者:
Yu, Kai
Yu, Kai
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ji, Shuiwang;Xu, Wei;Yu, Kai

文献摘要

被引文献

相似文献

我们考虑了监控视频中人类行为的自动识别。目前大多数方法基于从原始输入计算的复杂手工特征构建分类器。卷积神经网络(cnn)是一种可以直接作用于原始输入的深度模型。然而,这些模型目前仅限于处理2D输入。在本文中,我们开发了一种新的用于动作识别的三维CNN模型。该模型通过三维卷积从空间和时间两个维度提取特征,从而捕获编码在多个相邻帧中的运动信息。所开发的模型从输入帧中生成多个信息通道,最终的特征表示将所有通道的信息组合在一起。为了进一步提高性能,我们提出用高级特征对输出进行正则化,并将各种不同模型的预测结合起来。我们将开发的模型应用于识别机场监控视频的真实环境中的人类行为,与基线方法相比,它们取得了更好的性能。
We consider the automated recognition of human actions in surveillance videos. Most current methods build classifiers based on complex handcrafted features computed from the raw inputs. Convolutional neural networks (CNNs) are a type of deep model that can act directly on the raw inputs. However, such models are currently limited to handling 2D inputs. In this paper, we develop a novel 3D CNN model for action recognition. This model extracts features from both the spatial and the temporal dimensions by performing 3D convolutions, thereby capturing the motion information encoded in multiple adjacent frames. The developed model generates multiple channels of information from the input frames, and the final feature representation combines information from all channels. To further boost the performance, we propose regularizing the outputs with high-level features and combining the predictions of a variety of different models. We apply the developed models to recognize human actions in the real-world environment of airport surveillance videos, and they achieve superior performance in comparison to baseline methods.