Monocular Depth Estimation Using Cues Inspired by Biological Vision Systems

Monocular Depth Estimation Using Cues Inspired by Biological Vision Systems
复制标题

DOI:
10.1109/icpr56361.2022.9956454
复制
发表时间:
2022-04
期刊:
2022 26th International Conference on Pattern Recognition (ICPR)
影响因子:
--
通讯作者:
Dylan Auty;K. Mikolajczyk
Dylan Auty;K. Mikolajczyk
中科院分区:
其他
文献类型:
--
作者:
Dylan Auty;K. Mikolajczyk

文献摘要

相似文献

单目深度估计(MDE)旨在将场景的RGB图像从同一相机视图转换为逐像素深度图。由于缺少信息,它从根本上是不适定的:任何单个图像都可以从许多可能的3D场景中获取。因此,MDE任务的一部分是学习图像中的哪些视觉线索可以用于深度估计,以及如何使用。由于训练数据受到注释成本的限制或网络容量受到计算能力的限制,这是具有挑战性的。在这项工作中,我们证明了显式地将视觉提示信息注入到模型中有利于深度估计。生物视觉系统的研究,我们专注于语义信息和先验知识的对象的大小和它们的关系,模仿的相对大小,熟悉的大小,和绝对大小的生物线索。我们使用最先进的语义和实例分割模型来提供外部信息,并利用语言嵌入来编码类之间的关系信息。我们还提供了一个先验的平均现实世界中的对象的大小。这种外部信息克服了数据可用性的限制,并确保给定网络的有限容量集中在已知有用的线索上,从而提高性能。我们通过实验验证了我们的假设,并在广泛使用的NYUD 2室内深度估计基准上评估了所提出的模型。结果表明,当语义信息、尺寸先验和实例尺寸与RGB图像一起明确提供时,深度预测得到了改进,并且我们的方法可以很容易地适应任何深度估计系统。沿着。
Monocular depth estimation (MDE) aims to transform an RGB image of a scene into a pixelwise depth map from the same camera view. It is fundamentally ill-posed due to missing information: any single image can have been taken from many possible 3D scenes. Part of the MDE task is, therefore, to learn which visual cues in the image can be used for depth estimation, and how. With training data limited by cost of annotation or network capacity limited by computational power, this is challenging.In this work we demonstrate that explicitly injecting visual cue information into the model is beneficial for depth estimation. Following research into biological vision systems, we focus on semantic information and prior knowledge of object sizes and their relations, to emulate the biological cues of relative size, familiar size, and absolute size. We use state-of-the-art semantic and instance segmentation models to provide external information, and exploit language embeddings to encode relational information between classes. We also provide a prior on the average real-world size of objects. This external information overcomes the limitation in data availability, and ensures that the limited capacity of a given network is focused on known-helpful cues, therefore improving performance. We experimentally validate our hypothesis and evaluate the proposed model on the widely used NYUD2 indoor depth estimation benchmark. The results show improvements in depth prediction when the semantic information, size prior and instance size are explicitly provided along with the RGB images, and our method can be easily adapted to any depth estimation system.