Holistic 3D Human and Scene Mesh Estimation from Single View Images

Holistic 3D Human and Scene Mesh Estimation from Single View Images
复制标题

DOI:
10.1109/cvpr46437.2021.00040
复制
发表时间:
2020-12
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Zhenzhen Weng;Serena Yeung
Zhenzhen Weng;Serena Yeung
中科院分区:
其他
文献类型:
--
作者:
Zhenzhen Weng;Serena Yeung

文献摘要

相似文献

3D世界限制了人体姿势,而人体姿势传达了周围物体的信息。事实上,从放置在室内场景中的人的单个图像,我们作为人类,善于通过我们对物理定律的了解以及对合理物体和人体姿势的先验感知来解决人体姿势和房间布局的模糊性。然而,很少有计算机视觉模型充分利用这一事实。在这项工作中,我们提出了一种整体可训练的模型,该模型可以从单个 RGB 图像中感知 3D 场景,估计相机姿势和房间布局,并重建人体和物体网格。通过对估计的各个方面施加一组全面且复杂的损失,我们表明我们的模型优于现有的人体网格方法和室内场景重建方法。据我们所知,这是第一个在网格级别输出物体和人类预测,并对场景和人类姿势进行联合优化的模型。
The 3D world limits the human body pose and the human body pose conveys information about the surrounding objects. Indeed, from a single image of a person placed in an indoor scene, we as humans are adept at resolving ambiguities of the human pose and room layout through our knowledge of the physical laws and prior perception of the plausible object and human poses. However, few computer vision models fully leverage this fact. In this work, we pro-pose a holistically trainable model that perceives the 3D scene from a single RGB image, estimates the camera pose and the room layout, and reconstructs both human body and object meshes. By imposing a set of comprehensive and sophisticated losses on all aspects of the estimations, we show that our model outperforms existing human body mesh methods and indoor scene reconstruction methods. To the best of our knowledge, this is the first model that outputs both object and human predictions at the mesh level, and performs joint optimization on the scene and human poses.