3-d depth reconstruction from a single still image

3-d depth reconstruction from a single still image
复制标题

DOI:
10.1007/s11263-007-0071-y
复制
发表时间:
2008-01-01
影响因子:
19.5
通讯作者:
Ng, Andrew Y.
Ng, Andrew Y.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Saxena, Ashutosh;Chung, Sung H.;Ng, Andrew Y.

文献摘要

被引文献

相似文献

我们认为,从一个单一的静态图像的三维深度估计的任务。我们采取监督学习方法来解决这个问题,其中我们开始收集单目图像的训练集(非结构化室内和室外环境,包括森林,人行道,树木,建筑物等)。以及它们对应的地面实况深度图然后,我们应用监督学习来预测深度图的值作为图像的函数。深度估计是一个具有挑战性的问题,因为局部特征本身不足以估计一个点的深度,并且需要考虑图像的全局上下文。我们的模型使用一个分层的,多尺度马尔可夫随机场(MRF),结合多尺度的本地和全球的图像功能,并在图像中的不同点的深度和深度之间的关系模型。我们表明,即使在非结构化的场景中,我们的算法往往能够恢复相当准确的深度图。我们进一步提出了一个模型,结合单眼线索和立体(三角测量)线索,以获得更准确的深度估计比可能单独使用单眼或立体提示。
We consider the task of 3-d depth estimation from a single still image. We take a supervised learning approach to this problem, in which we begin by collecting a training set of monocular images (of unstructured indoor and outdoor environments which include forests, sidewalks, trees, buildings, etc.) and their corresponding ground-truth depthmaps. Then, we apply supervised learning to predict the value of the depthmap as a function of the image. Depth estimation is a challenging problem, since local features alone are insufficient to estimate depth at a point, and one needs to consider the global context of the image. Our model uses a hierarchical, multiscale Markov Random Field (MRF) that incorporates multiscale local- and global-image features, and models the depths and the relation between depths at different points in the image. We show that, even on unstructured scenes, our algorithm is frequently able to recover fairly accurate depthmaps. We further propose a model that incorporates both monocular cues and stereo (triangulation) cues, to obtain significantly more accurate depth estimates than is possible using either monocular or stereo cues alone.