Geometry-Free View Synthesis: Transformers and no 3D Priors

Geometry-Free View Synthesis: Transformers and no 3D Priors
复制标题

DOI:
10.1109/iccv48922.2021.01409
复制
发表时间:
2021-04
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Robin Rombach;Patrick Esser;B. Ommer
Robin Rombach;Patrick Esser;B. Ommer
中科院分区:
其他
文献类型:
--
作者:
Robin Rombach;Patrick Esser;B. Ommer

文献摘要

相似文献

从一张图像合成新的观点需要几何模型吗?由于受到局部卷积的约束,CNN需要显式的3D偏差来模拟几何变换。相反,我们演示了基于转换器的模型可以合成全新的视图,而不存在任何手工设计的3D偏差。这是通过以下方式实现的:(I)用于隐含地学习源和目标视图之间的远程3D对应的全局注意机制,以及(Ii)捕捉从单个图像预测新视图所固有的模糊性所必需的概率公式,从而克服了以前的方法被限制为相对较小的视点变化的限制。我们评估了将3D先验数据集成到变压器架构中的各种方法。然而,我们的实验表明,该变换不需要这样的几何先验知识,并且能够隐含地学习图像之间的3D关系。此外,这种方法在视觉质量方面优于最先进的技术,同时覆盖了可能实现的全部分布。
Is a geometric model required to synthesize novel views from a single image? Being bound to local convolutions, CNNs need explicit 3D biases to model geometric transformations. In contrast, we demonstrate that a transformer-based model can synthesize entirely novel views without any hand-engineered 3D biases. This is achieved by (i) a global attention mechanism for implicitly learning long-range 3D correspondences between source and target views, and (ii) a probabilistic formulation necessary to capture the ambiguity inherent in predicting novel views from a single image, thereby overcoming the limitations of previous approaches that are restricted to relatively small viewpoint changes. We evaluate various ways to integrate 3D priors into a transformer architecture. However, our experiments show that no such geometric priors are required and that the transformer is capable of implicitly learning 3D relationships between images. Furthermore, this approach outperforms the state of the art in terms of visual quality while covering the full distribution of possible realizations.