Benefits of Synthetically Pre-trained Depth-Prediction Networks for Indoor/Outdoor Image Classification

Benefits of Synthetically Pre-trained Depth-Prediction Networks for Indoor/Outdoor Image Classification
复制标题

DOI:
10.1109/wacvw58289.2023.00040
复制
发表时间:
2023-01
期刊:
2023 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW)
影响因子:
--
通讯作者:
Ke Lin;Irene Cho;Ameya S. Walimbe;Bryan A. Zamora;Alex Rich;Sirius Z. Zhang;Tobias Höllerer
Ke Lin;Irene Cho;Ameya S. Walimbe;Bryan A. Zamora;Alex Rich;Sirius Z. Zhang;Tobias Höllerer
中科院分区:
其他
文献类型:
--
作者:
Ke Lin;Irene Cho;Ameya S. Walimbe;Bryan A. Zamora;Alex Rich;Sirius Z. Zhang;Tobias Höllerer

文献摘要

相似文献

地面真实深度信息对于许多计算机视觉任务来说是必要的。收集这些信息具有挑战性,尤其是对于室外场景。在这项工作中,我们建议利用在合成场景上预训练的单视图深度预测神经网络来生成相对深度,我们称之为伪深度。这种方法是一种不太昂贵的选择,因为预先训练的神经网络从合成场景中获取准确的深度信息,不需要任何昂贵的传感器设备并且花费的时间更少。我们通过训练有或没有伪深度的室内/室外二元分类器来测量预训练神经网络的伪深度的有用性。我们还比较了使用伪深度和地面真实深度之间的准确性差异。我们通过实验表明,在训练中添加伪深度比 DIODE(大型标准测试数据集)上的非深度基线模型实现了 4.4% 的性能提升,保留了在 RGB 和地面真实深度上训练分类器所实现的 63.8% 的性能提升。它还将另一个数据集 SUN397 的性能提高了 1.3%,该数据集的地面实况深度不可用。我们的结果表明,可以从在合成场景上预先训练的模型中获取信息,并成功地将其应用到合成领域之外的现实世界数据中。
Ground truth depth information is necessary for many computer vision tasks. Collecting this information is chal-lenging, especially for outdoor scenes. In this work, we propose utilizing single-view depth prediction neural networks pre-trained on synthetic scenes to generate relative depth, which we call pseudo-depth. This approach is a less expen-sive option as the pre-trained neural network obtains ac-curate depth information from synthetic scenes, which does not require any expensive sensor equipment and takes less time. We measure the usefulness of pseudo-depth from pre-trained neural networks by training indoor/outdoor binary classifiers with and without it. We also compare the difference in accuracy between using pseudo-depth and ground truth depth. We experimentally show that adding pseudo-depth to training achieves a 4.4% performance boost over the non-depth baseline model on DIODE, a large stan-dard test dataset, retaining 63.8% of the performance boost achieved from training a classifier on RGB and ground truth depth. It also boosts performance by 1.3% on another dataset, SUN397, for which ground truth depth is not avail-able. Our result shows that it is possible to take information obtained from a model pre-trained on synthetic scenes and successfully apply it beyond the synthetic domain to real-world data.