Learning to Render Novel Views from Wide-Baseline Stereo Pairs

Learning to Render Novel Views from Wide-Baseline Stereo Pairs
复制标题

DOI:
10.1109/cvpr52729.2023.00481
复制
发表时间:
2023-04
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yilun Du;Cameron Smith;A. Tewari;V. Sitzmann
Yilun Du;Cameron Smith;A. Tewari;V. Sitzmann
中科院分区:
其他
文献类型:
--
作者:
Yilun Du;Cameron Smith;A. Tewari;V. Sitzmann

文献摘要

相似文献

我们介绍了一种新的视图合成方法,只给出了一个宽基线立体图像对。在这种具有挑战性的制度,3D场景点定期观察只有一次,需要基于先验的重建场景的几何形状和外观。我们发现,现有的方法,以新的视图合成稀疏观测失败,由于恢复不正确的3D几何形状,并由于高成本的微分渲染,排除了他们的规模大规模的训练。我们采取了一个步骤来解决这些缺点,制定了多视图Transformer编码器,提出了一个有效的,图像空间极线采样方案,以组装图像功能的目标射线,和一个轻量级的交叉注意力为基础的渲染器。我们的贡献使我们的方法能够在室内和室外场景的大规模真实世界数据集上进行训练。我们证明,我们的方法学习强大的多视图几何先验,同时减少渲染时间。我们在两个真实世界的数据集上进行了广泛的比较,显着优于以前的工作,从稀疏图像观测的新视图合成,并实现多视图一致的新视图合成。
We introduce a method for novel view synthesis given only a single wide-baseline stereo image pair. In this challenging regime, 3D scene points are regularly observed only once, requiring prior-based reconstruction of scene geometry and appearance. We find that existing approaches to novel view synthesis from sparse observations fail due to recovering incorrect 3D geometry and due to the high cost of differentiable rendering that precludes their scaling to large-scale training. We take a step towards resolving these shortcomings by formulating a multi-view transformer encoder, proposing an efficient, image-space epipolar line sampling scheme to assemble image features for a target ray, and a lightweight cross-attention-based renderer. Our contributions enable training of our method on a large-scale real-world dataset of indoor and outdoor scenes. We demonstrate that our method learns powerful multi-view geometry priors while reducing the rendering time. We conduct extensive comparisons on held-out test scenes across two real-world datasets, significantly outperforming prior work on novel view synthesis from sparse image observations and achieving multi-viewconsistent novel view synthesis.