Neural Groundplans: Persistent Neural Scene Representations from a Single Image

Neural Groundplans: Persistent Neural Scene Representations from a Single Image
复制标题

DOI:
--
复制
发表时间:
2022-07
期刊:
--
影响因子:
--
通讯作者:
Prafull Sharma;A. Tewari;Yilun Du;Sergey Zakharov;Rares Ambrus;Adrien Gaidon;W. Freeman;F. Durand;J. Tenenbaum;V. Sitzmann
Prafull Sharma;A. Tewari;Yilun Du;Sergey Zakharov;Rares Ambrus;Adrien Gaidon;W. Freeman;F. Durand;J. Tenenbaum;V. Sitzmann
中科院分区:
其他
文献类型:
--
作者:
Prafull Sharma;A. Tewari;Yilun Du;Sergey Zakharov;Rares Ambrus;Adrien Gaidon;W. Freeman;F. Durand;J. Tenenbaum;V. Sitzmann

文献摘要

相似文献

我们提出了一种将场景的 2D 图像观察映射到持久 3D 场景表示的方法,从而实现场景的可移动和不可移动组件的新颖视图合成和解开表示。受视觉和机器人技术中常用的鸟瞰图(BEV)表示的启发,我们提出了条件神经平面图、地面对齐的 2D 特征网格,作为持久且内存高效的场景表示。我们的方法是使用可微渲染从未标记的多视图观察中进行自我监督训练,并学习完成遮挡区域的几何形状和外观。此外,我们还表明,我们可以在训练时利用多视图视频来学习在测试时从单个图像中分别重建场景的静态和可移动组件。单独重建可移动对象的能力可以使用简单的启发式方法实现各种下游任务,例如提取以对象为中心的 3D 表示、新颖的视图合成、实例级分割、3D 边界框预测和场景编辑。这凸显了神经平面作为高效 3D 场景理解模型支柱的价值。
We present a method to map 2D image observations of a scene to a persistent 3D scene representation, enabling novel view synthesis and disentangled representation of the movable and immovable components of the scene. Motivated by the bird's-eye-view (BEV) representation commonly used in vision and robotics, we propose conditional neural groundplans, ground-aligned 2D feature grids, as persistent and memory-efficient scene representations. Our method is trained self-supervised from unlabeled multi-view observations using differentiable rendering, and learns to complete geometry and appearance of occluded regions. In addition, we show that we can leverage multi-view videos at training time to learn to separately reconstruct static and movable components of the scene from a single image at test time. The ability to separately reconstruct movable objects enables a variety of downstream tasks using simple heuristics, such as extraction of object-centric 3D representations, novel view synthesis, instance-level segmentation, 3D bounding box prediction, and scene editing. This highlights the value of neural groundplans as a backbone for efficient 3D scene understanding models.