Skeletonmae: Spatial-Temporal Masked Autoencoders for Self-Supervised Skeleton Action Recognition

Skeletonmae: Spatial-Temporal Masked Autoencoders for Self-Supervised Skeleton Action Recognition
复制标题

DOI:
10.1109/icmew59549.2023.00045
复制
发表时间:
2022-09
期刊:
2023 IEEE International Conference on Multimedia and Expo Workshops (ICMEW)
影响因子:
--
通讯作者:
Wenhan Wu;Yilei Hua;Ce Zheng;Shi-Bao Wu;Chen Chen-Chen;Aidong Lu
Wenhan Wu;Yilei Hua;Ce Zheng;Shi-Bao Wu;Chen Chen-Chen;Aidong Lu
中科院分区:
其他
文献类型:
--
作者:
Wenhan Wu;Yilei Hua;Ce Zheng;Shi-Bao Wu;Chen Chen-Chen;Aidong Lu

文献摘要

相似文献

近年来,基于自监督的人脸动作识别受到了越来越多的关注。通过利用未标记的数据,可以学习更多的泛化特征,以减轻过拟合问题,并减少对大量标记训练数据的需求。受MAE [1]的启发,我们提出了一个时空掩蔽自动编码器框架,用于自监督3D基于动作的动作识别(MAEetonMAE)。遵循MAE的掩蔽和重构流水线,我们利用基于卷积的编码器-解码器Transformer架构来重构掩蔽的骨架序列。提出了一种新的骨架序列联合级和帧级的时空掩蔽策略。这种预训练策略使得编码器输出具有空间和时间依赖性的可概括的骨架特征。给定未掩蔽的骨架序列,编码器针对动作识别任务进行微调。大量的实验表明,我们的EQUETONMAE实现了显着的性能,并优于国家的最先进的方法在NTU RGB+D 60和NTU RGB+D 120数据集。
Self-supervised skeleton-based action recognition has attracted more attention in recent years. By utilizing the unlabeled data, more generalizable features can be learned to alleviate the overfitting problem and reduce the demand for massive labeled training data. Inspired by the MAE [1], we propose a spatial-temporal masked autoencoder framework for self-supervised 3D skeleton-based action recognition (SkeletonMAE). Following MAE's masking and reconstruction pipeline, we utilize a skeleton-based encoder-decoder transformer architecture to reconstruct the masked skeleton sequences. A novel masking strategy, named Spatial-Temporal Masking, is introduced in terms of both joint-level and frame-level for the skeleton sequence. This pre-training strategy makes the encoder output generalizable skeleton features with spatial and temporal dependencies. Given the unmasked skeleton sequence, the encoder is fine-tuned for the action recognition task. Extensive experiments show that our SkeletonMAE achieves remarkable performance and outperforms the state-of-the-art methods on both NTU RGB+D 60 and NTU RGB+D 120 datasets.