D$^2$NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video

D$^2$NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video
复制标题

DOI:
--
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Tianhao Wu;Fangcheng Zhong;A. Tagliasacchi;Forrester Cole;Cengiz Oztireli
Tianhao Wu;Fangcheng Zhong;A. Tagliasacchi;Forrester Cole;Cengiz Oztireli
中科院分区:
其他
文献类型:
--
作者:
Tianhao Wu;Fangcheng Zhong;A. Tagliasacchi;Forrester Cole;Cengiz Oztireli

文献摘要

相似文献

给定单目视频,在恢复静态环境的同时分割解耦动态对象是机器智能中一个被广泛研究的问题。现有的解决方案通常在图像域处理这个问题,限制了它们的性能和对环境的理解。我们引入了解耦动态神经辐射场(D$^2$NeRF),这是一种自监督方法,它采用单目视频并学习3D场景表示,将移动物体(包括它们的阴影)与静态背景解耦。我们的方法通过两个独立的神经辐射场来表示运动物体和静态背景,其中只有一个允许时间变化。这种方法的幼稚实现会导致动态组件取代静态组件,因为前者的表示本质上更一般,容易过度拟合。为此,我们提出了一种新的损失方法来促进现象的正确分离。我们进一步提出了一个阴影场网络来检测和解耦动态移动的阴影。我们引入了一个包含各种动态对象和阴影的新数据集,并证明我们的方法在动态和静态3D对象解耦、遮挡和阴影去除以及移动对象的图像分割方面可以比最先进的方法获得更好的性能。
Given a monocular video, segmenting and decoupling dynamic objects while recovering the static environment is a widely studied problem in machine intelligence. Existing solutions usually approach this problem in the image domain, limiting their performance and understanding of the environment. We introduce Decoupled Dynamic Neural Radiance Field (D$^2$NeRF), a self-supervised approach that takes a monocular video and learns a 3D scene representation which decouples moving objects, including their shadows, from the static background. Our method represents the moving objects and the static background by two separate neural radiance fields with only one allowing for temporal changes. A naive implementation of this approach leads to the dynamic component taking over the static one as the representation of the former is inherently more general and prone to overfitting. To this end, we propose a novel loss to promote correct separation of phenomena. We further propose a shadow field network to detect and decouple dynamically moving shadows. We introduce a new dataset containing various dynamic objects and shadows and demonstrate that our method can achieve better performance than state-of-the-art approaches in decoupling dynamic and static 3D objects, occlusion and shadow removal, and image segmentation for moving objects.