IRISformer: Dense Vision Transformers for Single-Image Inverse Rendering in Indoor Scenes

IRISformer: Dense Vision Transformers for Single-Image Inverse Rendering in Indoor Scenes
复制标题

DOI:
10.1109/cvpr52688.2022.00284
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Rui Zhu;Zhengqin Li;J. Matai;F. Porikli;Manmohan Chandraker
Rui Zhu;Zhengqin Li;J. Matai;F. Porikli;Manmohan Chandraker
中科院分区:
其他
文献类型:
--
作者:
Rui Zhu;Zhengqin Li;J. Matai;F. Porikli;Manmohan Chandraker

文献摘要

相似文献

由于任意不同的物体形状、空间变化的材料和复杂的照明之间的无数相互作用,室内场景表现出显着的外观变化。由可见和不可见光源引起的阴影、高光和相互反射需要对逆渲染的远程交互进行推理,其目的是恢复图像形成的组成部分,即形状、材质和照明。在这项工作中,我们的直觉是,变压器架构学习到的远程注意力非常适合解决单图像逆渲染中长期存在的挑战。我们通过密集视觉转换器 IRISformer 的具体实例进行了演示,它在逆渲染所需的单任务和多任务推理方面表现出色。具体来说,我们提出了一种变压器架构,可以根据室内场景的单个图像同时估计深度、法线、空间变化的反照率、粗糙度和照明。我们对基准数据集的广泛评估展示了上述每项任务的最先进结果,使诸如对象插入和材质编辑之类的应用程序能够在单个不受约束的真实图像中实现,并且比以前的作品具有更高的真实感。代码和数据公开发布。11https://github.com/ViLab-UCSD/IRISformer
Indoor scenes exhibit significant appearance variations due to myriad interactions between arbitrarily diverse object shapes, spatially-changing materials, and complex lighting. Shadows, highlights, and inter-reflections caused by visible and invisible light sources require reasoning about long-range interactions for inverse rendering, which seeks to recover the components of image formation, namely, shape, material, and lighting. In this work, our intuition is that the long-range attention learned by transformer architectures is ideally suited to solve longstanding challenges in single-image inverse rendering. We demonstrate with a specific instantiation of a dense vision transformer, IRISformer, that excels at both single-task and multi-task reasoning required for inverse rendering. Specifically, we propose a transformer architecture to simultaneously estimate depths, normals, spatially-varying albedo, roughness and lighting from a single image of an indoor scene. Our extensive evaluations on benchmark datasets demonstrate state-of-the-art results on each of the above tasks, enabling applications like object insertion and material editing in a single unconstrained real image, with greater photorealism than prior works. Code and data are publicly released.11https://github.com/ViLab-UCSD/IRISformer