Neural Volumes: Learning Dynamic Renderable Volumes from images

Neural Volumes: Learning Dynamic Renderable Volumes from images
复制标题

DOI:
10.1145/3306346.3323020
复制
发表时间:
2019-07-01
影响因子:
6.2
通讯作者:
Sheikh, Yaser
Sheikh, Yaser
中科院分区:
计算机科学1区
文献类型:
--
作者:
Lombardi, Stephen;Simon, Tomas;Sheikh, Yaser

文献摘要

被引文献

相似文献

动态场景的建模和渲染具有挑战性,因为自然场景通常包含复杂的现象,例如薄结构、不断演变的拓扑、半透明、散射、遮挡和生物运动。在这些情况下,基于网格的重建和跟踪通常会失败,而其他方法(例如光场视频)通常依赖于受限的观看条件,这限制了交互性。我们通过提出一种基于学习的方法来表示动态对象,其灵感来自于断层扫描成像中使用的积分投影模型,从而克服了这些困难。该方法直接从多视图捕获设置中的 2D 图像进行监督,不需要显式重建或跟踪对象。我们的方法有两个主要组件:将输入图像转换为 3D 体积表示的编码器-解码器网络,以及支持端到端训练的可微光线行进操作。凭借其 3D 表示,与屏幕空间渲染技术相比,我们的构造可以更好地推断新的视点。编码器-解码器架构学习动态场景的潜在表示,使我们能够生成训练期间未见过的新颖内容序列。为了克服基于体素表示的内存限制,我们学习了在光线行进期间使用扭曲场实现的动态不规则网格结构。这种结构极大地提高了表观分辨率并减少了网格状伪影和锯齿状运动。最后,我们以面部表现捕捉为例,演示如何将基于表面的表示合并到我们的体积学习框架中,以适应需要最高分辨率的应用。
Modeling and rendering of dynamic scenes is challenging, as natural scenes often contain complex phenomena such as thin structures, evolving topology, translucency, scattering, occlusion, and biological motion. Mesh-based reconstruction and tracking often fail in these cases, and other approaches (e.g., light field video) typically rely on constrained viewing conditions, which limit interactivity. We circumvent these difficulties by presenting a learning-based approach to representing dynamic objects inspired by the integral projection model used in tomographic imaging. The approach is supervised directly from 2D images in a multi-view capture setting and does not require explicit reconstruction or tracking of the object. Our method has two primary components: an encoder-decoder network that transforms input images into a 3D volume representation, and a differentiable ray-marching operation that enables end-to-end training. By virtue of its 3D representation, our construction extrapolates better to novel viewpoints compared to screen-space rendering techniques. The encoder-decoder architecture learns a latent representation of a dynamic scene that enables us to produce novel content sequences not seen during training. To overcome memory limitations of voxel-based representations, we learn a dynamic irregular grid structure implemented with a warp field during ray-marching. This structure greatly improves the apparent resolution and reduces grid-like artifacts and jagged motion. Finally, we demonstrate how to incorporate surface-based representations into our volumetric-learning framework for applications where the highest resolution is required, using facial performance capture as a case in point.