Revealing Occlusions with 4D Neural Fields

Revealing Occlusions with 4D Neural Fields
复制标题

DOI:
10.1109/cvpr52688.2022.00302
复制
发表时间:
2022-04
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Basile Van Hoorick;Purva Tendulkar;Dídac Surís;Dennis Park;Simon Stent;Carl Vondrick
Basile Van Hoorick;Purva Tendulkar;Dídac Surís;Dennis Park;Simon Stent;Carl Vondrick
中科院分区:
其他
文献类型:
--
作者:
Basile Van Hoorick;Purva Tendulkar;Dídac Surís;Dennis Park;Simon Stent;Carl Vondrick

文献摘要

被引文献

相似文献

为了让计算机视觉系统在动态环境中运行,它们需要能够表示和推理对象的永久性。我们介绍了一个学习从单目RGB-D视频中估计4D视觉表示的框架,该框架能够持久保存对象,即使它们被遮挡了。与传统的视频表示不同,我们将点云编码为连续表示,这允许模型跨时空上下文参与解决遮挡问题。在与本文一起发布的两个大型视频数据集上,我们的实验表明,该表示能够在不改变任何体系结构的情况下,成功地揭示几个任务的遮挡。可视化显示,注意机制会自动学习跟随被遮挡的对象。由于我们的方法可以端到端地训练,并且很容易适应,我们相信它将有助于处理许多视频理解任务中的遮挡。数据、代码和模型随时可用。CS.澳大利亚。艾杜。
For computer vision systems to operate in dynamic situations, they need to be able to represent and reason about object permanence. We introduce a framework for learning to estimate 4D visual representations from monocular RGB-D video, which is able to persist objects, even once they become obstructed by occlusions. Unlike traditional video representations, we encode point clouds into a continuous representation, which permits the model to attend across the spatiotemporal context to resolve occlusions. On two large video datasets that we release along with this paper, our experiments show that the representation is able to successfully reveal occlusions for several tasks, without any architectural changes. Visualizations show that the attention mechanism automatically learns to follow occluded objects. Since our approach can be trained end-to-end and is easily adaptable, we believe it will be useful for handling occlusions in many video understanding tasks. Data, code, and models are available at occ1usions. cs. co1umbia. edu.