Mixed Neural Voxels for Fast Multi-view Video Synthesis

Mixed Neural Voxels for Fast Multi-view Video Synthesis
复制标题

用于快速多视图视频合成的混合神经体素

DOI:
--
复制
发表时间:
2022
期刊:
IEEE International Conference on Computer Vision
影响因子:
--
通讯作者:
Huaping Liu
Huaping Liu
中科院分区:
--
文献类型:
--
作者:
Feng Wang;Sinan Tan;Xinghang Li;Zeyue Tian;Huaping Liu

文献摘要

参考文献

被引文献

相似文献

由于现实世界环境和高动态运动的复杂性,从现实世界的多视图输入合成高保真视频具有挑战性。先前基于神经辐射场的工作已经证明了动态场景的高质量重建。然而,在现实场景中训练此类模型非常耗时,通常需要数天或数周的时间。在本文中,我们提出了一种名为 MixVoxels 的新颖方法来有效地表示动态场景,从而实现快速训练和渲染速度。所提出的 MixVoxels 将 4D 动态场景表示为静态和动态体素的混合,并使用不同的网络对其进行处理。这样,静态体素所需模态的计算可以通过轻量级模型来处理,这从本质上减少了计算量,因为许多日常动态场景以静态背景为主。为了区分这两种体素,我们提出了一种新的变化场来估计每个体素的时间方差。对于动态表示,我们设计了一种内积时间查询方法来有效地查询多个时间步,这对于恢复高动态运动至关重要。结果,通过输入 300 帧视频的动态场景的 15 分钟训练,MixVoxels 实现了比以前的方法更好的 PSNR。对于渲染,MixVoxels 可以以 37 fps 渲染 1K 分辨率的新颖视图视频。代码和训练模型可在 https://github.com/fengres/mixvoxels 获取。
Synthesizing high-fidelity videos from real-world multi-view input is challenging due to the complexities of real-world environments and high-dynamic movements. Previous works based on neural radiance fields have demonstrated high-quality reconstructions of dynamic scenes. However, training such models on real-world scenes is time-consuming, usually taking days or weeks. In this paper, we present a novel method named MixVoxels to efficiently represent dynamic scenes, enabling fast training and rendering speed. The proposed MixVoxels represents the 4D dynamic scenes as a mixture of static and dynamic voxels and processes them with different networks. In this way, the computation of the required modalities for static voxels can be processed by a lightweight model, which essentially reduces the amount of computation as many daily dynamic scenes are dominated by static backgrounds. To distinguish the two kinds of voxels, we propose a novel variation field to estimate the temporal variance of each voxel. For the dynamic representations, we design an inner product time query method to efficiently query multiple time steps, which is essential to recover the high-dynamic movements. As a result, with 15 minutes of training for dynamic scenes with inputs of 300-frame videos, MixVoxels achieves better PSNR than previous methods. For rendering, MixVoxels can render a novel view video with 1K resolution at 37 fps. Codes and trained models are available at https://github.com/fengres/mixvoxels.
来自无约束多视图视频的动态事件的 4D 可视化
DOI: 10.1109/cvpr42600.2020.00541
发表时间: 2020
期刊: IEEE Conference on Computer Vision and Pattern Recognition
影响因子: --
作者:
Bansal, Aayush;Vo, Minh;Sheikh, Yaser;Ramanan, Deva;Narasimhan, Srinivasa
通讯作者: Narasimhan, Srinivasa