Rotationally-Consistent Novel View Synthesis for Humans

Rotationally-Consistent Novel View Synthesis for Humans
复制标题

DOI:
10.1145/3394171.3413754
复制
发表时间:
2020-10
期刊:
Proceedings of the 28th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
Youngjoon Kwon;Stefano Petrangeli;Dahun Kim;Haoliang Wang;H. Fuchs;Viswanathan Swaminathan
Youngjoon Kwon;Stefano Petrangeli;Dahun Kim;Haoliang Wang;H. Fuchs;Viswanathan Swaminathan
中科院分区:
其他
文献类型:
--
作者:
Youngjoon Kwon;Stefano Petrangeli;Dahun Kim;Haoliang Wang;H. Fuchs;Viswanathan Swaminathan

文献摘要

相似文献

人类新颖的观点综合旨在综合从一个或多个参考观点拍摄的输入图像的人类主题的目标视图。尽管无模型的新观点合成方面取得了重大进展,但现有方法在像人类这样的复杂形状上应用了两个主要局限性。首先,这些方法主要集中于简单和对称对象,例如汽车和椅子,将其性能限制在细粒度和不对称形状上。其次,现有方法无法保证同一对象的不同相邻视图的视觉一致性。为了解决这些问题,我们在本文中介绍了人类受试者的新观点综合的学习框架,该框架明确地跨越了该主题的不同生成的观点。具体而言,我们在学习过程中引入了一种新颖的多视图监督和明确的旋转损失,使模型能够保留详细的身体部位并在相邻的合成视图之间达到一致性。为了显示我们方法的卓越性能,我们对收集的多视图人类动作(MVHA)数据集提出了定性和定量结果(由使用不同的MoCap序列动画的3D人类模型组成,并从54个不同观点中捕获),姿势变化的人类模型(PVHM)数据集和Shapenet。定性和定量结果表明,我们的方法在均观看质量中均超过了最先进的基线,并且在各种场景中,在多种情况下,在多种相邻视图中,在各种情况下,对于人类和僵化的对象,都可以在多个相邻视图中保持旋转一致性和复杂形状(例如细粒度的细节,具有挑战性的姿势)。
Human novel view synthesis aims to synthesize target views of a human subject given input images taken from one or more reference viewpoints. Despite significant advances in model-free novel view synthesis, existing methods present two major limitations when applied to complex shapes like humans. First, these methods mainly focus on simple and symmetric objects, e.g., cars and chairs, limiting their performances to fine-grained and asymmetric shapes. Second, existing methods cannot guarantee visual consistency across different adjacent views of the same object. To solve these problems, we present in this paper a learning framework for the novel view synthesis of human subjects, which explicitly enforces consistency across different generated views of the subject. Specifically, we introduce a novel multi-view supervision and an explicit rotational loss during the learning process, enabling the model to preserve detailed body parts and to achieve consistency between adjacent synthesized views. To show the superior performance of our approach, we present qualitative and quantitative results on the Multi-View Human Action (MVHA) dataset we collected (consisting of 3D human models animated with different Mocap sequences and captured from 54 different viewpoints), the Pose-Varying Human Model (PVHM) dataset, and ShapeNet. The qualitative and quantitative results demonstrate that our approach outperforms the state-of-the-art baselines in both per-view synthesis quality, and in preserving rotational consistency and complex shapes (e.g. fine-grained details, challenging poses) across multiple adjacent views in a variety of scenarios, for both humans and rigid objects.