Camera Pose Matters: Improving Depth Prediction by Mitigating Pose Distribution Bias

Camera Pose Matters: Improving Depth Prediction by Mitigating Pose Distribution Bias
复制标题

DOI:
10.1109/cvpr46437.2021.01550
复制
发表时间:
2021-06
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yunhan Zhao;Shu Kong;Charless C. Fowlkes
Yunhan Zhao;Shu Kong;Charless C. Fowlkes
中科院分区:
其他
文献类型:
--
作者:
Yunhan Zhao;Shu Kong;Charless C. Fowlkes

文献摘要

被引文献

相似文献

单目深度预测器通常在大规模训练集上进行训练,这些训练集自然偏向于相机姿态的分布。因此,训练的预测器无法对在不常见的相机姿势下捕获的测试示例进行可靠的深度预测。为了解决这个问题,我们提出了两种新的技术,在训练和预测过程中利用相机的姿态。首先,我们介绍了一个简单的视角感知数据增强,通过以几何一致的方式扰动现有的视图,合成具有更多样化视图的新训练示例。其次,我们提出了一个条件模型,利用每图像相机构成的先验知识,将其编码为输入的一部分。我们表明,联合应用这两种方法可以提高对在不常见甚至从未见过的相机姿势下捕获的图像的深度预测。我们表明,我们的方法提高性能时,应用到一系列不同的预测架构。最后,我们表明,显式编码的相机姿态分布提高了综合训练的深度预测的泛化性能时,评价真实的图像。
Monocular depth predictors are typically trained on large-scale training sets which are naturally biased w.r.t the distribution of camera poses. As a result, trained predictors fail to make reliable depth predictions for testing examples captured under uncommon camera poses. To address this issue, we propose two novel techniques that exploit the camera pose during training and prediction. First, we introduce a simple perspective-aware data augmentation that synthesizes new training examples with more diverse views by perturbing the existing ones in a geometrically consistent manner. Second, we propose a conditional model that exploits the per-image camera pose as prior knowledge by encoding it as a part of the input. We show that jointly applying the two methods improves depth prediction on images captured under uncommon and even never-before-seen camera poses. We show that our methods improve performance when applied to a range of different predictor architectures. Lastly, we show that explicitly encoding the camera pose distribution improves the generalization performance of a synthetically trained depth predictor when evaluated on real images.