Rethinking the ST-GCNs for 3D skeleton-based human action recognition

Rethinking the ST-GCNs for 3D skeleton-based human action recognition
复制标题

DOI:
10.1016/j.neucom.2021.05.004
复制
发表时间:
2021-05
期刊:
影响因子:
6
通讯作者:
Wei Peng;Jingang Shi;Tuomas Varanka;Guoying Zhao
Wei Peng;Jingang Shi;Tuomas Varanka;Guoying Zhao
中科院分区:
计算机科学2区
文献类型:
--
作者:
Wei Peng;Jingang Shi;Tuomas Varanka;Guoying Zhao

文献摘要

被引文献

相似文献

骨骼数据已经成为人类动作识别任务的替代方案,因为与传统的RGB输入相比,它提供了更紧凑和独特的信息。然而,与RGB输入不同的是,骨架数据位于非欧几里德空间,传统的深度学习方法无法充分发挥其潜力。幸运的是,随着几何深度学习的兴起,人们提出了时空图卷积网络(ST-GCN)来处理骨骼数据中的动作识别问题。ST-GCN及其变体非常适合基于骨骼的动作识别,并正在成为这一任务的主流框架。然而,通过固定骨架关节相关性或提供计算代价高昂的策略来构建骨架的动态拓扑,阻碍了任务的效率和性能。我们认为,这些操作中的许多对任务来说要么是不必要的,要么是有害的。通过对最新的ST-GCNS进行理论和实验分析,我们提供了一种简单而有效的策略来捕获全局图相关性,从而有效地对输入图序列的表示进行建模。此外,全局图策略还将图序列简化到欧几里德空间,从而引入多尺度时态过滤器来有效地捕捉动态信息。该方法不仅能够用更少的参数更好地提取图的相关性(仅为当前最优方法的12.6%),而且取得了较好的性能。在当前最大的3D数据集NTU-RGB+D和NTU-RGB+D 120上的大量实验表明,我们的网络能够在这一任务上执行高效和轻量级的优先处理。
The skeletal data has been an alternative for the human action recognition task as it provides more compact and distinct information compared to the traditional RGB input. However, unlike the RGB input, the skeleton data lies in a non-Euclidean space that traditional deep learning methods are not able to use their fullest potential. Fortunately, with the emerging trend of Geometric deep learning, the spatial-temporal graph convolutional network (ST-GCN) has been proposed to deal with the action recognition problem from skeleton data. ST-GCN and its variants fit well with skeleton-based action recognition and are becoming the mainstream frameworks for this task. However, the efficiency and the performance of the task are hindered by either fixing the skeleton joint correlations or providing a computational expensive strategy to construct a dynamic topology for the skeleton. We argue that many of these operations are either unnecessary or even harmful for the task. By theoretically and experimentally analysing the state-of-the-art ST-GCNs, we provide a simple but efficient strategy to capture the global graph correlations and thus efficiently model the representation of the input graph sequences. Moreover, the global graph strategy also reduces the graph sequence into the Euclidean space, thus a multi-scale temporal filter is introduced to efficiently capture the dynamic information. With the method, we are not only able to better extract the graph correlations with much fewer parameters (only 12.6% of the current best), but we also achieve a superior performance. Extensive experiments on current largest 3D datasets, NTU-RGB+D and NTU-RGB+D 120, demonstrate the ability of our network to perform efficient and lightweight priority on this task.