Learning 3D Semantic Scene Graphs From 3D Indoor Reconstructions

Learning 3D Semantic Scene Graphs From 3D Indoor Reconstructions
复制标题

DOI:
10.1109/cvpr42600.2020.00402
复制
发表时间:
2020-04
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Johanna Wald;Helisa Dhamo;N. Navab;Federico Tombari
Johanna Wald;Helisa Dhamo;N. Navab;Federico Tombari
中科院分区:
其他
文献类型:
--
作者:
Johanna Wald;Helisa Dhamo;N. Navab;Federico Tombari

文献摘要

被引文献

相似文献

场景理解一直是计算机视觉领域的研究热点。它不仅包括识别场景中的对象,还包括它们在给定上下文中的关系。有了这个目标,最近的一系列工作解决了3D语义分割和场景布局预测。在我们的工作中,我们专注于场景图,一个数据结构,组织在一个图形中的场景的实体,其中对象的节点和它们的关系建模为边缘。我们利用场景图的推理作为一种方式来进行3D场景理解,映射对象及其关系。特别是,我们提出了一个学习的方法,从场景的点云回归场景图。我们的新架构是基于PointNet和图卷积网络(GCN)。此外,我们介绍了3DSSG,一个半自动生成的数据集,其中包含语义丰富的场景图的3D场景。我们展示了我们的方法在领域不可知的检索任务中的应用,其中图形作为3D-3D和2D-3D匹配的中间表示。
Scene understanding has been of high interest in computer vision. It encompasses not only identifying objects in a scene, but also their relationships within the given context. With this goal, a recent line of works tackles 3D semantic segmentation and scene layout prediction. In our work we focus on scene graphs, a data structure that organizes the entities of a scene in a graph, where objects are nodes and their relationships modeled as edges. We leverage inference on scene graphs as a way to carry out 3D scene understanding, mapping objects and their relationships. In particular, we propose a learned method that regresses a scene graph from the point cloud of a scene. Our novel architecture is based on PointNet and Graph Convolutional Networks (GCN). In addition, we introduce 3DSSG, a semiautomatically generated dataset, that contains semantically rich scene graphs of 3D scenes. We show the application of our method in a domain-agnostic retrieval task, where graphs serve as an intermediate representation for 3D-3D and 2D-3D matching.