HaLP: Hallucinating Latent Positives for Skeleton-based Self-Supervised Learning of Actions

HaLP: Hallucinating Latent Positives for Skeleton-based Self-Supervised Learning of Actions
复制标题

DOI:
10.1109/cvpr52729.2023.01807
复制
发表时间:
2023-04
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Anshul B. Shah;A. Roy;Ketul Shah;Shlok Kumar Mishra;David W. Jacobs;A. Cherian;Ramalingam Chellappa
Anshul B. Shah;A. Roy;Ketul Shah;Shlok Kumar Mishra;David W. Jacobs;A. Cherian;Ramalingam Chellappa
中科院分区:
其他
文献类型:
--
作者:
Anshul B. Shah;A. Roy;Ketul Shah;Shlok Kumar Mishra;David W. Jacobs;A. Cherian;Ramalingam Chellappa

文献摘要

相似文献

最近,用于动作识别的骨架序列编码器的监督学习受到了极大的关注。然而,学习这种没有标签的编码器仍然是一个具有挑战性的问题。虽然先前的工作已经通过将对比学习应用于姿势序列而显示出有希望的结果,但通常观察到学习的表示的质量与用于制作正面的数据增强密切相关。然而,增强姿势序列是一项困难的任务,因为需要强制执行骨骼关节之间的几何约束,以使增强对于该动作是真实的。在这项工作中,我们提出了一种新的对比学习方法来训练模型,用于无标签的基于动作的动作识别。我们的主要贡献是一个简单的模块,HaLP -幻觉潜在的积极对比学习。具体来说,HaLP探索潜在空间的姿势在合适的方向,以产生新的积极的。为此,我们提出了一种新的优化配方,以解决其硬度显式控制的合成阳性。我们提出了近似的目标,使他们在封闭的形式,以最小的开销可解。我们通过实验表明,在标准的对比学习框架中使用这些生成的积极因素,可以在诸如NTU-60、NTU-120和PKU-II等基准测试中,在线性评估、迁移学习和kNN评估等任务上实现一致的改进。我们的代码可以在https://github.com/anshulbshah/HaLP上找到。
Supervised learning of skeleton sequence encoders for action recognition has received significant attention in recent times. However, learning such encoders without labels continues to be a challenging problem. While prior works have shown promising results by applying contrastive learning to pose sequences, the quality of the learned representations is often observed to be closely tied to data augmentations that are used to craft the positives. However, augmenting pose sequences is a difficult task as the geometric constraints among the skeleton joints need to be enforced to make the augmentations realistic for that action. In this work, we propose a new contrastive learning approach to train models for skeleton-based action recognition without labels. Our key contribution is a simple module, HaLP - to Hallucinate Latent Positives for contrastive learning. Specifically, HaLP explores the latent space of poses in suitable directions to generate new positives. To this end, we present a novel optimization formulation to solve for the synthetic positives with an explicit control on their hardness. We propose approximations to the objective, making them solvable in closed form with minimal overhead. We show via experiments that using these generated positives within a standard contrastive learning framework leads to consistent improvements across benchmarks such as NTU-60, NTU-120, and PKU-II on tasks like linear evaluation, transfer learning, and kNN evaluation. Our code can be found at https://github.com/anshulbshah/HaLP.