Joint Depth and Defocus Estimation From a Single Image Using Physical Consistency

Joint Depth and Defocus Estimation From a Single Image Using Physical Consistency
复制标题

DOI:
10.1109/tip.2021.3061901
复制
发表时间:
2021-03
影响因子:
10.6
通讯作者:
Anmei Zhang;Jian Sun
Anmei Zhang;Jian Sun
中科院分区:
计算机科学1区
文献类型:
--
作者:
Anmei Zhang;Jian Sun

文献摘要

相似文献

深度估计和离焦图是计算机视觉中的两个基本任务。近年来,许多方法借助深度学习强大的特征学习能力分别探索这两个任务,并取得了令人瞩目的进展。然而,由于在真实图像上难以密集标记深度和离焦,这些方法大多基于合成训练数据集,学习后的网络在真实图像上的性能下降明显。在本文中,我们解决了一个新的任务,即从单个图像中联合估计深度和离焦。我们设计了一个具有两个子网的双网络,分别用于估计深度和离焦。该网络在具有物理约束的合成数据集上进行联合训练,以加强深度和离焦之间的物理一致性。此外,我们设计了一种简单的方法来标记真实图像数据集上的深度和离焦顺序,并设计了两个新的度量标准来衡量真实图像上的深度和离焦估计的精度。综合实验表明,利用物理一致性约束对深度和离焦估计进行联合训练,使这两个子网能够相互引导,有效地提高了它们在真实离焦图像数据集上的深度和离焦估计性能。
Estimating depth and defocus maps are two fundamental tasks in computer vision. Recently, many methods explore these two tasks separately with the help of the powerful feature learning ability of deep learning and these methods have achieved impressive progress. However, due to the difficulty in densely labeling depth and defocus on real images, these methods are mostly based on synthetic training dataset, and the performance of learned network degrades significantly on real images. In this paper, we tackle a new task that jointly estimates depth and defocus from a single image. We design a dual network with two subnets respectively for estimating depth and defocus. The network is jointly trained on synthetic dataset with a physical constraint to enforce the physical consistency between depth and defocus. Moreover, we design a simple method to label depth and defocus order on real image dataset, and design two novel metrics to measure accuracies of depth and defocus estimation on real images. Comprehensive experiments demonstrate that joint training for depth and defocus estimation using physical consistency constraint enables these two subnets to guide each other, and effectively improves their depth and defocus estimation performance on real defocused image dataset.