A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose Estimation

A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose Estimation
复制标题

DOI:
10.48550/arxiv.2311.03312
复制
发表时间:
2023-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Qi-jun Zhao;Ce Zheng;Mengyuan Liu;Chen Chen-Chen
Qi-jun Zhao;Ce Zheng;Mengyuan Liu;Chen Chen-Chen
中科院分区:
其他
文献类型:
--
作者:
Qi-jun Zhao;Ce Zheng;Mengyuan Liu;Chen Chen-Chen

文献摘要

相似文献

将2D姿态序列提升到3D的3D人体姿态估计中的主导范例严重依赖于长期时间线索(即,使用令人生畏的视频帧数量)来提高精度,这会导致性能饱和、难以处理的计算和非因果问题。这可以归因于它们固有的无法感知空间背景,因为普通的2D关节坐标没有视觉线索。为了解决这个问题,我们提出了一个简单而强大的解决方案:利用现成的(预先训练的)2D姿态检测器产生的现成的中间视觉表示-甚至不需要对3D任务进行微调。关键的观察结果是,虽然姿态检测器学习定位2D关节,但这种表示(例如,特征地图)由于骨干网络中的区域操作而隐式地编码以联合为中心的空间上下文。我们设计了一个简单的基线名为上下文感知的PoseFormer来展示其有效性。在不访问任何时间信息的情况下,所提出的方法在速度和精度方面明显优于其上下文不可知的对应物PoseFormer和使用多达数百个视频帧的其他最先进的方法。项目页面:https://qitaozhao.github.io/ContextAware-PoseFormer
The dominant paradigm in 3D human pose estimation that lifts a 2D pose sequence to 3D heavily relies on long-term temporal clues (i.e., using a daunting number of video frames) for improved accuracy, which incurs performance saturation, intractable computation and the non-causal problem. This can be attributed to their inherent inability to perceive spatial context as plain 2D joint coordinates carry no visual cues. To address this issue, we propose a straightforward yet powerful solution: leveraging the readily available intermediate visual representations produced by off-the-shelf (pre-trained) 2D pose detectors -- no finetuning on the 3D task is even needed. The key observation is that, while the pose detector learns to localize 2D joints, such representations (e.g., feature maps) implicitly encode the joint-centric spatial context thanks to the regional operations in backbone networks. We design a simple baseline named Context-Aware PoseFormer to showcase its effectiveness. Without access to any temporal information, the proposed method significantly outperforms its context-agnostic counterpart, PoseFormer, and other state-of-the-art methods using up to hundreds of video frames regarding both speed and precision. Project page: https://qitaozhao.github.io/ContextAware-PoseFormer