HybridAvatar: Efficient Mesh-based Human Avatar Generation from Few-Shot Monocular Images with Implicit Mesh Displacement

HybridAvatar: Efficient Mesh-based Human Avatar Generation from Few-Shot Monocular Images with Implicit Mesh Displacement
复制标题

DOI:
10.1109/ismar-adjunct60411.2023.00080
复制
发表时间:
2023-10
期刊:
2023 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct)
影响因子:
--
通讯作者:
Tianxing Fan;Bangbang Yang;Chong Bao;Lei Wang;Guofeng Zhang;Zhaopeng Cui
Tianxing Fan;Bangbang Yang;Chong Bao;Lei Wang;Guofeng Zhang;Zhaopeng Cui
中科院分区:
其他
文献类型:
--
作者:
Tianxing Fan;Bangbang Yang;Chong Bao;Lei Wang;Guofeng Zhang;Zhaopeng Cui

文献摘要

被引文献

相似文献

从商品级摄像机的少量拍摄图像中高效且可控地生成人体化身是AR/VR应用所需要的功能之一。然而,现有的方法要么依赖于像SMPL这样的参数模型,这种模型不能很好地拟合真实形状,也缺乏几何细节(如衣服上的皱纹),或者利用神经隐函数来恢复细节,但需要仔细的数据收集和密集的计算,这使得模型无法部署到现实世界的应用中。本文提出了一种新的基于单目RGB图像生成3D人体头像的框架--HyBridge Avtal.该框架不仅可以恢复细节几何形状,而且能够以较低的计算代价支持姿势动画.与以往的工作不同,我们的方法同时利用了显式参数模型和神经隐式函数的优点,通过学习数据驱动的隐式位移场来补充参数网格模型的细节。为了同时实现高保真建模和即时推理,我们设计了一种级联机制来对人体形状进行两步建模,并提出了一种基于球谐函数的可微纹理处理来编码人体外观。结合网格模型和基于纹理的绘制策略,实现了VR/AR应用中人体动画的快速绘制。实验表明,与SOTA方法相比,该方法在少镜头RGB图像下具有更好的三维人体建模性能,并且能够高效地制作人体化身动画。
Efficient and controllable human avatar generation from few-shot images of a commodity-level camera is one of the desired functions for AR/VR applications. However, existing methods either rely on a parametric model like SMPL, which cannot closely fit the real shape and also lacks geometric details (e.g., wrinkles in clothes), or utilize neural implicit functions to recover details while requiring careful data collection and intensive computation, which prohibits the model deployment into real-world applications. In this paper, we propose a novel framework, called HybridAvatar, for 3D human avatar generation from monocular RGB images, which not only recovers detailed geometries, but also supports pose animation with low-cost computation. Different from previous works, our method takes advantage of both the explicit parametric model and the neural implicit function, and learns a data-driven implicit displacement field to complement details upon the parametric mesh model. To achieve both high-fidelity modeling and instant inference, we design a cascaded mechanism to model body shapes in a two-stage manner and propose a spherical harmonics-based differentiable texturing process to encode human appearances. With the advantage of mesh model and texture-based rendering strategy, we achieve fast rendering of human animations in VR/AR applications. Experiments demonstrate the improved 3D human body modeling performance of our method over SOTA approaches under few-shot RGB images and the ability to animate human avatars efficiently.