Semantic MapNet: Building Allocentric SemanticMaps and Representations from Egocentric Views

Semantic MapNet: Building Allocentric SemanticMaps and Representations from Egocentric Views
复制标题

Semantic MapNet:从自我中心视图构建异中心语义地图和表示

DOI:
--
复制
发表时间:
2020
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Dhruv Batra
Dhruv Batra
中科院分区:
--
文献类型:
--
作者:
Vincent Cartillier;Zhile Ren;Neha Jain;Stefan Lee;Irfan Essa;Dhruv Batra

文献摘要

参考文献

被引文献

相似文献

我们研究语义映射的任务-具体来说,一个具体的代理(机器人或自我中心的AI助手)被赋予一个新的环境之旅,并要求建立一个allocentric自上而下的语义地图(“什么是在哪里?”)从具有已知姿态的RGB-D相机的自我中心观察(经由定位传感器)。重要的是,我们的目标是建立3D空间的神经情景记忆和空间语义表示,使智能体能够轻松地学习同一空间中的后续任务-导航到参观期间看到的对象(“查找椅子”)或回答有关空间的问题(“你在房子里看到了多少把椅子?”)。为了实现这一目标,我们提出了语义映射网(SMNet),它包括:(1)编码每个以自我为中心的RGB-D帧的以自我为中心的视觉编码器,(2)将以自我为中心的特征投影到平面图上的适当位置的特征投影器,(3)学习积累投影的以自我为中心的特征的尺寸为平面图长度×宽度×特征尺寸的空间记忆张量,以及(4)使用记忆张量来产生语义自上而下映射的映射解码器。SMNet结合了(已知的)投影相机几何和神经表征学习的优势。在Matterport 3D数据集中的语义映射任务中,SMNet在平均IoU指标上的表现明显优于竞争基线4.01 - 16.81%(绝对值),在边界F1指标上的表现则为3.81 - 19.69%(绝对值)。此外,我们展示了如何使用SMNet构建的空间语义allocentric表示的任务ObjectNav的和隐藏的问题检索。项目页面:https://vincentcartillier.github.io/smnet.html。
We study the task of semantic mapping – specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentric top-down semantic map (‘what is where?’) from egocentric observations of an RGB-D camera with known pose (via localization sensors). Importantly, our goal is to build neural episodic memories and spatio-semantic representations of 3D spaces that enable the agent to easily learn subsequent tasks in the same space – navigating to objects seen during the tour (‘Find chair’) or answering questions about the space (‘How many chairs did you see in the house?’). Towards this goal, we present Semantic MapNet (SMNet), which consists of: (1) an Egocentric Visual Encoder that encodes each egocentric RGB-D frame, (2) a Feature Projector that projects egocentric features to appropriate locations on a floor-plan, (3) a Spatial Memory Tensor of size floor-plan length×width×feature-dims that learns to accumulate projected egocentric features, and (4) a Map Decoder that uses the memory tensor to produce semantic top-down maps. SMNet combines the strengths of (known) projective camera geometry and neural representation learning. On the task of semantic mapping in the Matterport3D dataset, SMNet significantly outperforms competitive baselines by 4.01−16.81% (absolute) on mean-IoU and 3.81−19.69% (absolute) on Boundary-F1 metrics. Moreover, we show how to use the spatio-semantic allocentric representations build by SMNet for the task of ObjectNav and Embodied Question Answering. Project page: https://vincentcartillier.github.io/smnet.html.
DOI: --
发表时间: 2018-10
期刊: ArXiv
影响因子: --
作者:
Valts Blukis;Dipendra Misra;Ross A. Knepper;Yoav Artzi
通讯作者: Valts Blukis;Dipendra Misra;Ross A. Knepper;Yoav Artzi