Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly- Throughs

Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly- Throughs
复制标题

DOI:
10.1109/cvpr52688.2022.01258
复制
发表时间:
2021-12
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Haithem Turki;Deva Ramanan;M. Satyanarayanan
Haithem Turki;Deva Ramanan;M. Satyanarayanan
中科院分区:
其他
文献类型:
--
作者:
Haithem Turki;Deva Ramanan;M. Satyanarayanan

文献摘要

相似文献

我们使用神经辐射场 (NeRF),通过主要从无人机收集的跨越建筑物甚至多个城市街区的大规模视觉捕捉来构建交互式 3D 环境。与单个对象场景(传统上评估 NeRF)相比,我们的规模带来了多种挑战,包括(1)需要对数千个具有不同光照条件的图像进行建模,每个图像仅捕获场景的一小部分,(2)模型容量过大,使得在单个 GPU 上进行训练变得不可行,以及(3)实现交互式飞行的快速渲染面临重大挑战。为了应对这些挑战,我们首先分析大规模场景的可见性统计数据,激发稀疏网络结构,其中参数专门针对场景的不同区域。我们引入了一种简单的数据并行几何聚类算法,该算法将训练图像(或更确切地说像素)划分为可以并行训练的不同 NeRF 子模块。我们在现有数据集(Quad 6k 和 UrbanScene3D)以及我们自己的无人机镜头上评估了我们的方法,将训练速度提高了 3 倍,PSNR 提高了 12%。我们还评估了 Mega-NeRF 之上的最新 NeRF 快速渲染器,并引入了一种利用时间相干性的新颖方法。我们的技术比传统 NeRF 渲染速度提高了 40 倍,同时 PSNR 质量保持在 0.8 db 以内,超过了现有快速渲染器的保真度。
We use neural radiance fields (NeRFs) to build interac-tive 3D environments from large-scale visual captures spanning buildings or even multiple city blocks collected pri-marily from drones. In contrast to single object scenes (on which NeRFs are traditionally evaluated), our scale poses multiple challenges including (1) the need to model thou-sands of images with varying lighting conditions, each of which capture only a small subset of the scene, (2) pro-hibitively large model capacities that make it infeasible to train on a single GPU, and (3) significant challenges for fast rendering that would enable interactive fly-throughs. To address these challenges, we begin by analyzing visi-bility statistics for large-scale scenes, motivating a sparse network structure where parameters are specialized to dif-ferent regions of the scene. We introduce a simple geomet-ric clustering algorithm for data parallelism that partitions training images (or rather pixels) into different NeRF sub-modules that can be trained in parallel. We evaluate our approach on existing datasets (Quad 6k and UrbanScene3D) as well as against our own drone footage, improving training speed by 3x and PSNR by 12%. We also evaluate re-cent NeRF fast renderers on top of Mega-NeRF and intro-duce a novel method that exploits temporal coherence. Our technique achieves a 40x speedup over conventional NeRF rendering while remaining within 0.8 db in PSNR quality, exceeding the fidelity of existing fast renderers.