HandyPose: Multi-level framework for hand pose estimation

HandyPose: Multi-level framework for hand pose estimation
复制标题

DOI:
10.1016/j.patcog.2022.108674
复制
发表时间:
2022-04-07
影响因子:
8
通讯作者:
Savakis, Andreas
Savakis, Andreas
中科院分区:
计算机科学1区
文献类型:
--
作者:
Gupta, Divyansh;Artacho, Bruno;Savakis, Andreas

文献摘要

被引文献

相似文献

由于手的自由度和关节的频繁遮挡,手的姿态估计是一项具有挑战性的任务。为了解决这些挑战,我们提出了HandyPose,这是一种使用单个RGB图像作为输入的用于2D手部姿势估计的单通道端到端可训练架构。该方法采用了一种具有多级特征的编码器-解码器框架,沿着采用了一种新颖的多级瀑布式空间池化模型来实现多尺度表示,在保持网络可管理的尺寸复杂度和模块性的同时,实现了手部姿态的高精度。HandyPose采用多尺度方法来表示上下文,通过在网络的各个级别上合并空间信息来减轻由于池化而导致的分辨率损失。我们的高级多级瀑布模块利用渐进级联过滤的效率,同时通过瀑布模块中不同网络级别的多级特征的级联来保持更大的视野。解码器结合了瀑布和多尺度功能,可在单个阶段生成准确的联合热图。我们的研究结果证明了流行的数据集上的最先进的性能,并表明HandyPose是一个强大而有效的2D手部姿势估计架构。(c)2022爱思唯尔有限公司保留所有权利。
Hand pose estimation is a challenging task due to the large number of degrees of freedom and the frequent occlusions of joints. To address these challenges, we propose HandyPose, a single-pass, end -to-end trainable architecture for 2D hand pose estimation using a single RGB image as input. Adopt-ing an encoder-decoder framework with multi-level features, along with a novel multi-level waterfall atrous spatial pooling module for multi-scale representations, our method achieves high accuracy in hand pose while maintaining manageable size complexity and modularity of the network. HandyPose takes a multi-scale approach to representing context by incorporating spatial information at various levels of the network to mitigate the loss of resolution due to pooling. Our advanced multi-level waterfall module leverages the efficiency of progressive cascade filtering while maintaining larger fields-of-view through the concatenation of multi-level features from different levels of the network in the waterfall module. The decoder incorporates both the waterfall and multi-scale features for the generation of accurate joint heatmaps in a single stage. Our results demonstrate state-of-the-art performance on popular datasets and show that HandyPose is a robust and efficient architecture for 2D hand pose estimation.(c) 2022 Elsevier Ltd. All rights reserved.