Deep Convolutional Neural Networks for Human Action Recognition Using Depth Maps and Postures

Deep Convolutional Neural Networks for Human Action Recognition Using Depth Maps and Postures
复制标题

使用深度图和姿势进行人类动作识别的深度卷积神经网络

DOI:
10.1109/tsmc.2018.2850149
复制
发表时间:
2019-09-01
影响因子:
8.7
通讯作者:
Feng, David Dagan
Feng, David Dagan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Kamel, Aouaidjia;Sheng, Bin;Feng, David Dagan

文献摘要

被引文献

相似文献

在本文中,我们提出了一种使用卷积神经网络(CNN)从深度图和姿势数据中识别人类动作的方法(D-Fusion)。两个输入描述符用于动作表示。第一输入是累积人类动作的连续深度图的深度运动图像,而第二输入是表示身体关节随时间的运动的提议的移动关节描述符。为了最大限度地提取特征以进行准确的动作分类,使用不同的输入训练三个CNN通道。第一通道用深度运动图像(DMI)训练,第二通道用DMI和移动关节描述符一起训练,并且第三通道仅用移动关节描述符训练。从三个CNN通道生成的动作预测被融合在一起,用于最终的动作分类。我们提出了几种融合评分操作,以最大限度地提高正确动作的得分。实验结果表明,将三个通道的输出进行融合,其效果优于单通道或两个通道的融合。我们所提出的方法在三个公共数据集上进行了评估:1)Microsoft action 3-D数据集(MSRD 3D); 2)德克萨斯大学达拉斯分校的多模态人类动作数据集; 3)多模态动作数据集(MAD)数据集。测试结果表明,该方法优于现有的大多数国家的最先进的方法,如直方图的4-D法向和MSRUNK 3D上的小波。虽然MAD数据集包含大量的动作(35个动作),但与现有的动作RGB-D数据集相比,本文超过了6.84%的最先进的方法。
In this paper, we present a method (Action-Fusion) for human action recognition from depth maps and posture data using convolutional neural networks (CNNs). Two input descriptors are used for action representation. The first input is a depth motion image that accumulates consecutive depth maps of a human action, whilst the second input is a proposed moving joints descriptor which represents the motion of body joints over time. In order to maximize feature extraction for accurate action classification, three CNN channels are trained with different inputs. The first channel is trained with depth motion images (DMIs), the second channel is trained with both DMIs and moving joint descriptors together, and the third channel is trained with moving joint descriptors only. The action predictions generated from the three CNN channels are fused together for the final action classification. We propose several fusion score operations to maximize the score of the right action. The experiments show that the results of fusing the output of three channels are better than using one channel or fusing two channels only. Our proposed method was evaluated on three public datasets: 1) Microsoft action 3-D dataset (MSRAction3D); 2) University of Texas at Dallas-multimodal human action dataset; and 3) multimodal action dataset (MAD) dataset. The testing results indicate that the proposed approach outperforms most of existing state-of-the-art methods, such as histogram of oriented 4-D normals and Actionlet on MSRAction3D. Although MAD dataset contains a high number of actions (35 actions) compared to existing action RGB-D datasets, this paper surpasses a state-of-the-art method on the dataset by 6.84%.