Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras

Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras
复制标题

DOI:
10.1109/tvcg.2018.2868527
复制
发表时间:
2018-09
影响因子:
5.2
通讯作者:
Young-Woon Cha;True Price;Zhen Wei;Xinran Lu;Nicholas Rewkowski;Rohan Chabra;Zihe Qin;Hyounghun Kim;Zhaoqi Su;Yebin Liu;A. Ilie;A. State;Zhenlin Xu;Jan-Michael Frahm;H. Fuchs
Young-Woon Cha;True Price;Zhen Wei;Xinran Lu;Nicholas Rewkowski;Rohan Chabra;Zihe Qin;Hyounghun Kim;Zhaoqi Su;Yebin Liu;A. Ilie;A. State;Zhenlin Xu;Jan-Michael Frahm;H. Fuchs
中科院分区:
计算机科学1区
文献类型:
--
作者:
Young-Woon Cha;True Price;Zhen Wei;Xinran Lu;Nicholas Rewkowski;Rohan Chabra;Zihe Qin;Hyounghun Kim;Zhaoqi Su;Yebin Liu;A. Ilie;A. State;Zhenlin Xu;Jan-Michael Frahm;H. Fuchs

文献摘要

被引文献

相似文献

我们提出了一种新的方法,在日常环境中的动态室内和室外场景的3D重建,仅利用用户佩戴的相机。这种方法允许在任何位置进行体验的3D重建和来自任何地方的虚拟图尔斯。所提出的以自我为中心的重建系统的关键创新是从近体视图(例如,用户眼镜上的相机)捕获佩戴者的身体姿势和面部表情,并且使用面向外的视图捕获周围环境。然而,以自我为中心的重建的主要挑战是近体视图的覆盖率差-也就是说,用户的身体和面部是从便于佩戴但不便于捕获的Vantage位置观察的。为了克服这些挑战,我们提出了一种基于参数模型的方法来进行用户运动估计。这种方法利用卷积神经网络(CNN)进行近视图身体姿势估计,我们引入了一种基于CNN的方法,用于结合音频和视频的面部表情估计。对于捕获期间的每个时间点,来自这些系统的中间基于模型的重建被用于重新瞄准用户的高保真预扫描模型。我们证明了所提出的自给自足的头戴式捕获系统能够在室内和室外情况下重建佩戴者的运动及其周围环境,而无需任何额外的视图。作为概念证明,我们展示了如何在虚拟现实系统中身临其境地体验所产生的3D加时间重建(例如,HTC Vive)。我们预计,所提出的以自我为中心的捕获和重建系统的大小最终将被减小,以适应未来的AR眼镜,并将广泛用于沉浸式3D远程呈现,虚拟图尔斯,和一般用途的任何地方的3D内容创建。
We propose a new approach for 3D reconstruction of dynamic indoor and outdoor scenes in everyday environments, leveraging only cameras worn by a user. This approach allows 3D reconstruction of experiences at any location and virtual tours from anywhere. The key innovation of the proposed ego-centric reconstruction system is to capture the wearer's body pose and facial expression from near-body views, e.g. cameras on the user's glasses, and to capture the surrounding environment using outward-facing views. The main challenge of the ego-centric reconstruction, however, is the poor coverage of the near-body views – that is, the user's body and face are observed from vantage points that are convenient for wear but inconvenient for capture. To overcome these challenges, we propose a parametric-model-based approach to user motion estimation. This approach utilizes convolutional neural networks (CNNs) for near-view body pose estimation, and we introduce a CNN-based approach for facial expression estimation that combines audio and video. For each time-point during capture, the intermediate model-based reconstructions from these systems are used to re-target a high-fidelity pre-scanned model of the user. We demonstrate that the proposed self-sufficient, head-worn capture system is capable of reconstructing the wearer's movements and their surrounding environment in both indoor and outdoor situations without any additional views. As a proof of concept, we show how the resulting 3D-plus-time reconstruction can be immersively experienced within a virtual reality system (e.g., the HTC Vive). We expect that the size of the proposed egocentric capture-and-reconstruction system will eventually be reduced to fit within future AR glasses, and will be widely useful for immersive 3D telepresence, virtual tours, and general use-anywhere 3D content creation.