Zero-Shot Depth Estimation from Light Field Using a Convolutional Neural Network

Zero-Shot Depth Estimation from Light Field Using a Convolutional Neural Network
复制标题

使用卷积神经网络从光场进行零样本深度估计

DOI:
10.1109/tci.2020.2967148
复制
发表时间:
2020
影响因子:
5.4
通讯作者:
Dong Liu
Dong Liu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jiayong Peng;Zhiwei Xiong;Yicheng Wang;Yueyi Zhang;Dong Liu

文献摘要

被引文献

相似文献

本文提出了一种基于零镜头学习的光场深度估计框架,该框架只学习从输入光场到相应视差图的端到端映射,不需要额外的训练数据,也不需要对真值深度进行监督。该方法克服了现有基于学习方法存在的两个主要困难,在实践中更加可行。首先,它省去了在训练过程中获取各种场景的底真深度作为标签的巨大负担。其次,它避免了应用于与训练数据截然不同的内容或在不同相机配置下捕获的光场时严重的域移位效应。另一方面,与传统的非学习方法相比,该方法更好地利用了四维光场的相关性,得到了更优的深度结果。此外,我们将这种零镜头学习框架扩展到光场视频的深度估计。我们首次证明,通过联合利用空间、角度和时间维度之间的相关性,可以从光场视频中估计出更准确、更稳健的深度。我们对合成光场图像数据集和真实光场图像数据集以及自采集的光场视频数据集进行了综合实验。定量和定性结果验证了我们的方法优于最先进的性能,特别是对于具有挑战性的现实世界场景。
This article proposes a zero-shot learning-based framework for light field depth estimation, which learns an end-to-end mapping solely from an input light field to the corresponding disparity map with neither extra training data nor supervision of groundtruth depth. The proposed method overcomes two major difficulties posed in existing learning-based methods and is thus much more feasible in practice. First, it saves the huge burden of obtaining groundtruth depth of a variety of scenes to serve as labels during training. Second, it avoids the severe domain shift effect when applied to light fields with drastically different content or captured under different camera configurations from the training data. On the other hand, compared with conventional non-learning-based methods, the proposed method better exploits the correlations in the 4D light field and generates much superior depth results. Moreover, we extend this zero-shot learning framework to depth estimation from light field videos. For the first time, we demonstrate that more accurate and robust depth can be estimated from light field videos by jointly exploiting the correlations across spatial, angular, and temporal dimensions. We conduct comprehensive experiments on both synthetic and real-world light field image datasets, as well as a self collected light field video dataset. Quantitative and qualitative results validate the superior performance of our method over the state-of-the-arts, especially for the challenging real-world scenes.