Geometric Deep Neural Network using Rigid and Non-Rigid Transformations for Human Action Recognition

Geometric Deep Neural Network using Rigid and Non-Rigid Transformations for Human Action Recognition
复制标题

DOI:
10.1109/iccv48922.2021.01238
复制
发表时间:
2021-10
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Rasha Friji;Hassen Drira;F. Chaieb;Hamza Kchok;S. Kurtek
Rasha Friji;Hassen Drira;F. Chaieb;Hamza Kchok;S. Kurtek
中科院分区:
其他
文献类型:
--
作者:
Rasha Friji;Hassen Drira;F. Chaieb;Hamza Kchok;S. Kurtek

文献摘要

相似文献

深度学习架构虽然在大多数计算机视觉任务中是成功的,但它是为具有底层欧几里得结构的数据而设计的,这通常无法实现,因为预处理的数据可能位于非线性空间上。在本文中,我们提出了一种几何感知的深度学习方法,使用刚性和非刚性变换优化来进行基于机器人的动作识别。骨架序列首先被建模为Kendall形状空间上的轨迹,然后映射到线性切空间。然后将得到的结构化数据馈送到深度学习架构,该架构包括一个优化3D骨架的刚性和非刚性转换的层,然后是CNN-LSTM网络。对两个大规模骨架数据集(即NTU-RGB+D和NTU-RGB+D 120)的评估已经证明,所提出的方法优于现有的几何深度学习方法,并且在大多数配置方面超过了最近发布的方法。
Deep Learning architectures, albeit successful in most computer vision tasks, were designed for data with an underlying Euclidean structure, which is not usually fulfilled since pre-processed data may lie on a non-linear space. In this paper, we propose a geometry aware deep learning approach using rigid and non rigid transformation optimization for skeleton-based action recognition. Skeleton sequences are first modeled as trajectories on Kendall’s shape space and then mapped to the linear tangent space. The resulting structured data are then fed to a deep learning architecture, which includes a layer that optimizes over rigid and non rigid transformations of the 3D skeletons, followed by a CNN-LSTM network. The assessment on two large scale skeleton datasets, namely NTU-RGB+D and NTU-RGB+D 120, has proven that the proposed approach outperforms existing geometric deep learning methods and exceeds recently published approaches with respect to the majority of configurations.