Semi-supervised Semantic Mapping Through Label Propagation with Semantic Texture Meshes

Semi-supervised Semantic Mapping Through Label Propagation with Semantic Texture Meshes
复制标题

DOI:
10.1007/s11263-019-01187-z
复制
发表时间:
2019-06
影响因子:
19.5
通讯作者:
R. Rosu;Jan Quenzel;Sven Behnke
R. Rosu;Jan Quenzel;Sven Behnke
中科院分区:
计算机科学2区
文献类型:
--
作者:
R. Rosu;Jan Quenzel;Sven Behnke

文献摘要

相似文献

场景理解是机器人在非结构化环境中的一项重要能力。虽然大多数SLAM方法提供场景的几何表示,但语义地图对于与周围环境的更复杂交互是必要的。当前的方法将语义图视为几何的一部分,这限制了可扩展性和准确性。我们建议表示的语义地图作为一个几何网格和语义纹理耦合在独立的分辨率。其关键思想是,在许多环境中的几何形状可以大大简化,而不会失去保真度,而语义信息可以存储在一个更高的分辨率,独立的网格。我们构建了一个网格深度传感器来表示场景的几何形状和融合信息到语义纹理从分割的各个RGB视图的场景。使语义持久化在一个全球性的网格,使我们能够执行时间和空间的一致性的个人视图预测。为此,我们提出了一种有效的方法,通过迭代地重新训练语义分割与存储在地图中的信息,并使用重新训练的分割,以拒绝融合的语义,建立个人分割之间的共识。我们证明了我们的方法的准确性和可扩展性重建语义地图的场景从NYUv2和场景跨越大型建筑物。
Scene understanding is an important capability for robots acting in unstructured environments. While most SLAM approaches provide a geometrical representation of the scene, a semantic map is necessary for more complex interactions with the surroundings. Current methods treat the semantic map as part of the geometry which limits scalability and accuracy. We propose to represent the semantic map as a geometrical mesh and a semantic texture coupled at independent resolution. The key idea is that in many environments the geometry can be greatly simplified without loosing fidelity, while semantic information can be stored at a higher resolution, independent of the mesh. We construct a mesh from depth sensors to represent the scene geometry and fuse information into the semantic texture from segmentations of individual RGB views of the scene. Making the semantics persistent in a global mesh enables us to enforce temporal and spatial consistency of the individual view predictions. For this, we propose an efficient method of establishing consensus between individual segmentations by iteratively retraining semantic segmentation with the information stored within the map and using the retrained segmentation to re-fuse the semantics. We demonstrate the accuracy and scalability of our approach by reconstructing semantic maps of scenes from NYUv2 and a scene spanning large buildings.