Single-View 3D Scene Reconstruction and Parsing by Attribute Grammar

Single-View 3D Scene Reconstruction and Parsing by Attribute Grammar
复制标题

DOI:
10.1109/tpami.2017.2689007
复制
发表时间:
2018-03-01
影响因子:
23.6
通讯作者:
Zhu, Song-Chun
Zhu, Song-Chun
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liu, Xiaobai;Zhao, Yibiao;Zhu, Song-Chun

文献摘要

被引文献

相似文献

在本文中,我们提出了一个属性语法,用于求解两个耦合任务:i)将2D图像解析为语义区域; ii)恢复所有地区的3D场景结构。所提出的语法由一组生产规则组成,每种制作规则都描述了3D场景中平面表面之间的一种空间关系。这些生产规则用于将输入图像分解为层次解析图表示,其中每个图节点表示平面表面或复合表面。与其他随机图像语法不同,所提出的语法增强每个图节点具有一组属性变量,以描绘场景级的全局几何形状,例如摄像机焦距或蜜蜂/几何形状,例如表面之间的表面正常,表面正常,表面之间的接触线。这些几何属性在解析图中施加了节点及其离子之间的约束。在概率框架下,我们开发了马尔可夫链蒙特卡洛方法来构建一个解析图,该图形同时优化了2D图像识别和3D场景重建目的。我们在公共基准和新收集的数据集上评估了我们的方法。实验表明,所提出的方法能够实现单个图像的最新场景重建。
In this paper, we present an attribute grammar for solving two coupled tasks: i) parsing a 2D image into semantic regions; and ii) recovering the 3D scene structures of all regions. The proposed grammar consists of a set of production rules, each describing a kind of spatial relation between planar surfaces in 3D scenes. These production rules are used to decompose an input image into a hierarchical parse graph representation where each graph node indicates a planar surface or a composite surface. Different from other stochastic image grammars, the proposed grammar augments each graph node with a set of attribute variables to depict scene-level global geometry, e.g., camera focal length, or bee/geometry, e.g., surface normal, contact lines between surfaces. These geometric attributes impose constraints between a node and its off-springs in the parse graph. Under a probabilistic framework, we develop a Markov Chain Monte Carlo method to construct a parse graph that optimizes the 2D image recognition and 3D scene reconstruction purposes simultaneously. We evaluated our method on both public benchmarks and newly collected datasets. Experiments demonstrate that the proposed method is capable of achieving state-of-the-art scene reconstruction of a single image.