Rotationally-Temporally Consistent Novel View Synthesis of Human Performance Video

Rotationally-Temporally Consistent Novel View Synthesis of Human Performance Video
复制标题

DOI:
10.1007/978-3-030-58548-8_23
复制
发表时间:
2020-08
期刊:
--
影响因子:
--
通讯作者:
Youngjoon Kwon;Stefano Petrangeli;Dahun Kim;Haoliang Wang;Eunbyung Park;Viswanathan Swaminathan;H. Fuchs
Youngjoon Kwon;Stefano Petrangeli;Dahun Kim;Haoliang Wang;Eunbyung Park;Viswanathan Swaminathan;H. Fuchs
中科院分区:
其他
文献类型:
--
作者:
Youngjoon Kwon;Stefano Petrangeli;Dahun Kim;Haoliang Wang;Eunbyung Park;Viswanathan Swaminathan;H. Fuchs

文献摘要

相似文献

新颖视点视频合成旨在合成新颖视点视频,给定从多个参考视点和连续时间步采集的人类表演的输入捕获。尽管在无模型新视图合成方面取得了很大的进步,但现有的方法在应用于复杂和时变的人类行为时存在三个局限性。首先,这些方法(和相关数据集)主要考虑简单和对称的对象。其次,它们不会在生成的视图之间强制显式的一致性。第三,他们关注静态和不移动的物体。因此,在跨不同视点或时间步合成时,人类主体的细粒度细节可能会出现不一致。为了应对这些挑战,我们引入了一个特定于人类的框架,该框架采用了学习的3d感知表示。具体来说,我们首先引入了一种新的暹罗网络,该网络采用门控层来更好地重建潜在的体积表示,从而获得最终的视觉结果。此外,连续时间步长的特征在网络内部共享,以提高时间一致性。其次,我们引入了一种新的损失来显式地强制在空间和时间上生成的视图的一致性。第三,我们提出了多视图人类行为(MVHA)数据集,包括从54个视点捕获的近1200个合成人类行为。在MVHA、Pose-Varying Human Model和ShapeNet数据集上的实验表明,我们的方法在视图生成质量和时空一致性方面都优于最先进的基线。
Novel viewvideosynthesis aims to synthesize novel viewpoints videos given input captures of a human performance taken from multiple reference viewpoints and over consecutive time steps. Despite great advances in model-free novel view synthesis, existing methods present three limitations when applied to complex and time-varying human performance. First, these methods (and related datasets) mainly consider simple and symmetric objects. Second, they do not enforce explicit consistency across generated views. Third, they focus on static and non-moving objects. The fine-grained details of a human subject can therefore suffer from inconsistencies when synthesized across different viewpoints or time steps. To tackle these challenges, we introduce a human-specific framework that employs a learned 3D-aware representation. Specifically, we first introduce a novel siamese network that employs a gating layer for better reconstruction of the latent volumetric representation and, consequently, final visual results. Moreover, features from consecutive time steps are shared inside the network to improve temporal consistency. Second, we introduce a novel loss to explicitly enforce consistency across generated views both inspaceand intime. Third, we present the Multi-View Human Action (MVHA) dataset, consisting of near 1200 synthetic human performance captured from 54 viewpoints. Experiments on the MVHA, Pose-Varying Human Model and ShapeNet datasets show that our method outperforms the state-of-the-art baselines both in view generation quality and spatio-temporal consistency.