DeFeat-Net: General Monocular Depth via Simultaneous Unsupervised Representation Learning

DeFeat-Net: General Monocular Depth via Simultaneous Unsupervised Representation Learning
复制标题

DOI:
10.1109/cvpr42600.2020.01441
复制
发表时间:
2020-03
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Jaime Spencer;R. Bowden;Simon Hadfield
Jaime Spencer;R. Bowden;Simon Hadfield
中科院分区:
其他
文献类型:
--
作者:
Jaime Spencer;R. Bowden;Simon Hadfield

文献摘要

相似文献

在目前的单目深度研究中,主要的方法是在大型数据集上采用无监督训练,由扭曲的光度一致性驱动。这种方法缺乏鲁棒性,并且无法推广到具有挑战性的领域,例如夜间场景或不利的天气条件,其中关于光度一致性的假设被打破。我们提出了深度和特征网络(Deep-Net),这是一种同时学习跨域密集特征表示的方法,以及基于扭曲特征一致性的鲁棒深度估计框架。所得到的特征表示是以无监督的方式学习的,不需要显式的地面实况对应。我们表明,在一个单一的域中,我们的技术与目前的单目深度估计和监督特征表示学习的最新技术水平相当。然而,通过同时学习特征、深度和运动,我们的技术能够推广到具有挑战性的领域,使DeepMin-Net在更具挑战性的序列(如夜间驾驶)上的所有误差测量减少约10%,从而超越当前最先进的技术。
In the current monocular depth research, the dominant approach is to employ unsupervised training on large datasets, driven by warped photometric consistency. Such approaches lack robustness and are unable to generalize to challenging domains such as nighttime scenes or adverse weather conditions where assumptions about photometric consistency break down. We propose DeFeat-Net (Depth & Feature network), an approach to simultaneously learn a cross-domain dense feature representation, alongside a robust depth-estimation framework based on warped feature consistency. The resulting feature representation is learned in an unsupervised manner with no explicit ground-truth correspondences required. We show that within a single domain, our technique is comparable to both the current state of the art in monocular depth estimation and supervised feature representation learning. However, by simultaneously learning features, depth and motion, our technique is able to generalize to challenging domains, allowing DeFeat-Net to outperform the current state-of-the-art with around 10% reduction in all error measures on more challenging sequences such as nighttime driving.