Action recognition using kinematics posture feature on 3D skeleton joint locations

Action recognition using kinematics posture feature on 3D skeleton joint locations
复制标题

DOI:
10.1016/j.patrec.2021.02.013
复制
发表时间:
2021-03-13
影响因子:
5.1
通讯作者:
Yagi, Yasushi
Yagi, Yasushi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ahad, Md Atiqur Rahman;Ahmed, Masud;Yagi, Yasushi

文献摘要

被引文献

相似文献

动作识别是计算机视觉及相关领域中一个非常广泛的研究领域。为了提高动作识别的性能,我们提出了基于骨架数据的三维关节位置的运动学位置特征(KPF)提取。在这种方法中,我们认为骨架三维关节作为运动学传感器。我们提出了线性关节位置特征(LJPF)和角度关节位置特征(AJPF)的基础上的三维线性关节的位置和骨段之间的角度。然后,我们将每个动作的每个视频帧的这两个运动学特征联合收割机组合起来,以创建KPF特征集。这些特征集在时域中对运动的变化进行编码,就好像每个身体关节表示运动学位置和取向传感器一样。在下一阶段,我们使用低通滤波器处理提取的KPF特征描述符,并使用具有优化长度的滑动窗口分割它们。该概念类似于处理运动学传感器数据的方法。从分割的窗口,我们计算基于位置的统计特征(PSF)。这些特征包括时域统计特征(例如,平均值、标准偏差、方差等)。这些统计特征编码姿势的变化(即,关节位置和角度)。为了进行分类,我们探索了支持向量机(线性),RNN,CNNRNN和ConvRNN模型。所提出的PSF特征集在基于统计机器学习和基于深度学习的模型中表现出出色的性能。为了进行评估,我们探索了五个基准数据集,即UTKinect-Round 3D,Kinect Activity Recognition Dataset(KARD),MSR 3D Action Pairs,佛罗伦萨3D和Office Activity Dataset(OAD)。为了防止过度拟合,我们考虑将leave-one-subject-out框架作为实验设置并执行10倍交叉验证。我们的方法在这些基准数据集上优于现有的几种方法,并取得了非常有前途的分类性能。(c)2021爱思唯尔有限公司版权所有。
Action recognition is a very widely explored research area in computer vision and related fields. We propose Kinematics Posture Feature (KPF) extraction from 3D joint positions based on skeleton data for improving the performance of action recognition. In this approach, we consider the skeleton 3D joints as kinematics sensors. We propose Linear Joint Position Feature (LJPF) and Angular Joint Position Feature (AJPF) based on 3D linear joint positions and angles between bone segments. We then combine these two kinematics features for each video frame for each action to create the KPF feature sets. These feature sets encode the variation of motion in the temporal domain as if each body joint represents kinematics position and orientation sensors. In the next stage, we process the extracted KPF feature descriptor by using a low pass filter, and segment them by using sliding windows with optimized length. This concept resembles the approach of processing kinematics sensor data. From the segmented windows, we compute the Position-based Statistical Feature (PSF). These features consist of temporal domain statistical features (e.g., mean, standard deviation, variance, etc.). These statistical features encode the variation of postures (i.e., joint positions and angles) across the video frames. For performing classification, we explore Support Vector Machine (Linear), RNN, CNNRNN, and ConvRNN model. The proposed PSF feature sets demonstrate prominent performance in both statistical machine learning-and deep learning-based models. For evaluation, we explore five benchmark datasets namely UTKinect-Action3D, Kinect Activity Recognition Dataset (KARD), MSR 3D Action Pairs, Florence 3D, and Office Activity Dataset (OAD). To prevent overfitting, we consider the leave-one-subject-out framework as the experimental setup and perform 10-fold cross-validation. Our approach outperforms several existing methods in these benchmark datasets and achieves very promising classification performance.(c) 2021 Elsevier B.V. All rights reserved.