MEBOW: Monocular Estimation of Body Orientation in the Wild

MEBOW: Monocular Estimation of Body Orientation in the Wild
复制标题

DOI:
10.1109/cvpr42600.2020.00351
复制
发表时间:
2020-06
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Chenyan Wu;Yukun Chen;Jiajia Luo;Che-Chun Su;A. Dawane;Bikramjot Hanzra;Zhuo Deng;Bilan Liu;J. Z. Wang;Cheng-Hao Kuo
Chenyan Wu;Yukun Chen;Jiajia Luo;Che-Chun Su;A. Dawane;Bikramjot Hanzra;Zhuo Deng;Bilan Liu;J. Z. Wang;Cheng-Hao Kuo
中科院分区:
其他
文献类型:
--
作者:
Chenyan Wu;Yukun Chen;Jiajia Luo;Che-Chun Su;A. Dawane;Bikramjot Hanzra;Zhuo Deng;Bilan Liu;J. Z. Wang;Cheng-Hao Kuo

文献摘要

相似文献

身体方向估计在许多应用中提供了关键的视觉线索,包括机器人和自动驾驶。当3-D姿态估计由于图像分辨率差、遮挡或不可区分的身体部位而难以推断时,这是特别期望的。我们提出了COCO-MEBOW(野外身体方向的单眼估计),这是一个新的大规模数据集,用于根据单个野外图像进行方向估计。COCO数据集的55 K图像中约130 K人体的身体方向标签已使用高效且高精度的注释管道收集。我们还验证了数据集的好处。首先,我们证明了我们的数据集可以大大提高人体方向估计模型的性能和鲁棒性,该模型的发展以前受到可用训练数据的规模和多样性的限制。此外,我们提出了一种新的三维人体姿态估计的三源解决方案,其中3-D姿态标签,2-D姿态标签和我们的身体方向标签都用于联合训练。我们的模型在单目3-D人体姿势估计方面明显优于最先进的双源解决方案,其中训练仅使用3-D姿势标签和2-D姿势标签。这证实了MEBOW用于3-D人体姿态估计的重要优势,这是特别有吸引力的,因为身体方向的每个实例标记成本远远低于3-D姿态。这项工作展示了MEBOW在解决涉及理解人类行为的现实挑战方面的巨大潜力。关于这项工作的进一步信息可在https://chenyanwu.github.io/MEBOW/上查阅。
Body orientation estimation provides crucial visual cues in many applications, including robotics and autonomous driving. It is particularly desirable when 3-D pose estimation is difficult to infer due to poor image resolution, occlusion or indistinguishable body parts. We present COCO-MEBOW (Monocular Estimation of Body Orientation in the Wild), a new large-scale dataset for orientation estimation from a single in-the-wild image. The body-orientation labels for around 130K human bodies within 55K images from the COCO dataset have been collected using an efficient and high-precision annotation pipeline. We also validated the benefits of the dataset. First, we show that our dataset can substantially improve the performance and the robustness of a human body orientation estimation model, the development of which was previously limited by the scale and diversity of the available training data. Additionally, we present a novel triple-source solution for 3-D human pose estimation, where 3-D pose labels, 2-D pose labels, and our body-orientation labels are all used in joint training. Our model significantly outperforms state-of-the-art dual-source solutions for monocular 3-D human pose estimation, where training only uses 3-D pose labels and 2-D pose labels. This substantiates an important advantage of MEBOW for 3-D human pose estimation, which is particularly appealing because the per-instance labeling cost for body orientations is far less than that for 3-D poses. The work demonstrates high potential of MEBOW in addressing real-world challenges involving understanding human behaviors. Further information of this work is available at https://chenyanwu.github.io/MEBOW/.