Action recognition based on joint trajectory maps with convolutional neural networks

Action recognition based on joint trajectory maps with convolutional neural networks
复制标题

基于卷积神经网络联合轨迹图的动作识别

DOI:
10.1016/j.knosys.2018.05.029
复制
发表时间:
2018-10-15
影响因子:
8.8
通讯作者:
Hou, Yonghong
Hou, Yonghong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wang, Pichao;Li, Wanqing;Hou, Yonghong

文献摘要

被引文献

相似文献

卷积神经网络(ConvNets)近年来在许多计算机视觉任务中表现出了良好的性能,尤其是基于图像的识别。如何有效地将转换网应用于基于序列的数据仍然是一个悬而未决的问题。提出了一种将三维骨骼序列中携带的时空信息表示为三幅二维图像的简单有效的方法,该方法将关节轨迹及其动力学编码为图像中的颜色分布,称为联合轨迹图(JTM),并采用ConvNets学习用于人体动作识别的鉴别特征。这种基于图像的表示使我们能够微调现有的用于骨架序列分类的ConvNets模型,而无需重新训练网络。这三个JTM在三个正交平面中生成,并相互提供互补信息。通过三个JTM的乘法分数融合,进一步提高了最终的识别率。该方法在NTU RGB+D大型数据集、MSRC-12 Kinect手势数据集(MSRC-12)、G3D数据集和UTD多模式人体行为数据集(UTD-MHAD)上进行了测试,取得了最好的结果。
Convolutional Neural Networks (ConvNets) have recently shown promising performance in many computer vision tasks, especially image-based recognition. How to effectively apply ConvNets to sequence-based data is still an open problem. This paper proposes an effective yet simple method to represent spatio-temporal information carried in 3D skeleton sequences into three 2D images by encoding the joint trajectories and their dynamics into color distribution in the images, referred to as Joint Trajectory Maps (JTM), and adopts ConvNets to learn the discriminative features for human action recognition. Such an image-based representation enables us to fine-tune existing ConvNets models for the classification of skeleton sequences without training the networks afresh. The three JTMs are generated in three orthogonal planes and provide complimentary information to each other. The final recognition is further improved through multiplicative score fusion of the three JTMs. The proposed method was evaluated on four public benchmark datasets, the large NTU RGB + D Dataset, MSRC-12 Kinect Gesture Dataset (MSRC-12), G3D Dataset and UTD Multimodal Human Action Dataset (UTD-MHAD) and achieved the state-of-the-art results.