3D-based Deep Convolutional Neural Network for action recognition with depth sequences

3D-based Deep Convolutional Neural Network for action recognition with depth sequences
复制标题

DOI:
10.1016/j.imavis.2016.04.004
复制
发表时间:
2016-11
期刊:
Image Vis. Comput.
影响因子:
--
通讯作者:
Zhi Liu;Chenyang Zhang;Yingli Tian
Zhi Liu;Chenyang Zhang;Yingli Tian
中科院分区:
其他
文献类型:
--
作者:
Zhi Liu;Chenyang Zhang;Yingli Tian

文献摘要

被引文献

相似文献

在过去的十年中,传统的手工设计动作识别特征的算法一直是一个热门的研究领域。与RGB视频相比,深度序列对光照变化更不敏感,并且由于其能够捕获对象的几何信息而更具鉴别力。与许多现有的依赖于精心设计的特征的动作识别方法不同,本文研究了使用深度序列和相应的骨骼关节信息的基于深度学习的动作识别。首先,我们构建了一个基于3D的深度卷积神经网络(3D 2CNN)来直接从原始深度序列中学习时空特征,然后通过考虑骨架关节之间的简单位置和角度信息,为每个序列计算一个基于关节的特征向量JointVector。最后,将3D 2CNN学习特征的支持向量机(SVM)分类结果与JointVector融合进行动作识别。实验结果表明,该方法可以从深度序列中学习到时不变和视点不变的特征表示。所提出的方法实现了与UTKinect-RISK 3D数据集上的最先进方法相当的结果,并且与MSR-RISK 3D数据集上的基线方法相比实现了上级性能。我们通过将学习到的特征从一个数据集(MSR-RIS 3D)转移到另一个数据集(UTKinect-RIS 3D)而无需再训练来进一步研究训练模型的泛化,并获得非常有希望的分类精度。
Traditional algorithms to design hand-crafted features for action recognition have been a hot research area in the last decade. Compared to RGB video, depth sequence is more insensitive to lighting changes and more discriminative due to its capability to catch geometric information of object. Unlike many existing methods for action recognition which depend on well-designed features, this paper studies deep learning-based action recognition using depth sequences and the corresponding skeleton joint information. Firstly, we construct a 3D-based Deep Convolutional Neural Network (3D2CNN) to directly learn spatio-temporal features from raw depth sequences, then compute a joint based feature vector named JointVector for each sequence by taking into account the simple position and angle information between skeleton joints. Finally, support vector machine (SVM) classification results from 3D2CNN learned features and JointVector are fused to take action recognition. Experimental results demonstrate that our method can learn feature representation which is time-invariant and viewpoint-invariant from depth sequences. The proposed method achieves comparable results to the state-of-the-art methods on the UTKinect-Action3D dataset and achieves superior performance in comparison to baseline methods on the MSR-Action3D dataset. We further investigate the generalization of the trained model by transferring the learned features from one dataset (MSR-Action3D) to another dataset (UTKinect-Action3D) without retraining and obtain very promising classification accuracy.