Learning Local-Global Contextual Adaptation for Multi-Person Pose Estimation

Learning Local-Global Contextual Adaptation for Multi-Person Pose Estimation
复制标题

DOI:
10.1109/cvpr52688.2022.01272
复制
发表时间:
2021-09
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Nan Xue;Tianfu Wu;Gui-Song Xia;L. Zhang
Nan Xue;Tianfu Wu;Gui-Song Xia;L. Zhang
中科院分区:
其他
文献类型:
--
作者:
Nan Xue;Tianfu Wu;Gui-Song Xia;L. Zhang

文献摘要

被引文献

相似文献

本文研究了自下而上的多人姿态估计问题。基于中心偏移量公式的局部化问题可以在理想情况下在局部窗口搜索方案中得到纠正这一新的强有力的观察结果,我们提出了一种多人姿态估计方法,称为LOGO-CAP,通过学习对人类姿态的局部-全局上下文适应。具体地说,该方法首先从局部小窗口中的局部关键点扩展图学习关键点吸引图,然后将其作为聚焦关键点的全局热图上的动态卷积核进行上下文自适应,从而实现准确的多人姿势估计。我们的方法是端到端可训练的,在单次前向传递中具有接近实时的推理速度,在用于自下而上的人体姿势估计的CoCo KeyPoint基准上获得了最先进的性能。使用COCO训练的模型,我们的方法在具有挑战性的OCHuman数据集上也远远超过了现有技术。
This paper studies the problem of multi-person pose estimation in a bottom-up fashion. With a new and strong observation that the localization issue of the center-offset formulation can be remedied in a local-window search scheme in an ideal situation, we propose a multi-person pose estimation approach, dubbed as LOGO-CAP, by learning the LOcal-GlObal Contextual Adaptation for human Pose. Specifically, our approach learns the keypoint attraction maps (KAMs) from the local keypoints expansion maps (KEMs) in small local windows in the first step, which are subsequently treated as dynamic convolutional kernels on the keypoints-focused global heatmaps for contextual adaptation, achieving accurate multi-person pose estimation. Our method is end-to-end trainable with near real-time inference speed in a single forward pass, obtaining state-of-the-art performance on the COCO keypoint benchmark for bottom-up human pose estimation. With the COCO trained model, our method also outperforms prior arts by a large margin on the challenging OCHuman dataset.