Spatial Temporal Graph Deconvolutional Network for Skeleton-Based Human Action Recognition

Spatial Temporal Graph Deconvolutional Network for Skeleton-Based Human Action Recognition
复制标题

DOI:
10.1109/lsp.2021.3049691
复制
发表时间:
2021-01-01
影响因子:
3.9
通讯作者:
Zhao, Guoying
Zhao, Guoying
中科院分区:
工程技术2区
文献类型:
--
作者:
Peng, Wei;Shi, Jingang;Zhao, Guoying

文献摘要

被引文献

相似文献

由于时空图卷积网络(ST-GCN)的强大能力,基于时空图卷积网络的人体动作识别已经取得了很大的成功。然而,通过消息传播的节点交互并不总是提供补充信息。相反,它甚至可能产生破坏性的噪音,从而使学习的表征无法区分。不可避免地,图形表示也会变得过于平滑,特别是当多个GCN层堆叠时。本文提出了时空图去卷积网络(ST-GDNs),一种新的和灵活的图形去卷积技术,以缓解这个问题。在其核心,该方法提供了一个更好的消息聚合,通过删除嵌入冗余的输入图无论是从节点,帧或元素的方式在不同的网络层。在三个当前最具挑战性的基准测试上进行的大量实验验证了ST-GDN在这些数据集上持续提高了性能并大大减小了模型大小。
Benefited from the powerful ability of spatial temporal Graph Convolutional Networks (ST-GCNs), skeleton-based human action recognition has gained promising success. However, the node interaction through message propagation does not always provide complementary information. Instead, it May even produce destructive noise and thus make learned representations indistinguishable. Inevitably, the graph representation would also become over-smoothing especially when multiple GCN layers are stacked. This paper proposes spatial-temporal graph deconvolutional networks (ST-GDNs), a novel and flexible graph deconvolution technique, to alleviate this issue. At its core, this method provides a better message aggregation by removing the embedding redundancy of the input graphs from either node-wise, frame-wise or element-wise at different network layers. Extensive experiments on three current most challenging benchmarks verify that ST-GDN consistently improves the performance and largely reduce the model size on these datasets.