Visual Representations for Semantic Target Driven Navigation

Visual Representations for Semantic Target Driven Navigation
复制标题

DOI:
10.1109/icra.2019.8793493
复制
发表时间:
2018-05
期刊:
2019 International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Arsalan Mousavian;Alexander Toshev;Marek Fiser;J. Kosecka;James Davidson
Arsalan Mousavian;Alexander Toshev;Marek Fiser;J. Kosecka;James Davidson
中科院分区:
其他
文献类型:
--
作者:
Arsalan Mousavian;Alexander Toshev;Marek Fiser;J. Kosecka;James Davidson

文献摘要

被引文献

相似文献

导航的良好视觉表现是什么?我们在语义视觉导航的背景下研究这个问题,这是一个机器人通过以前看不见的环境找到目标物体的问题,例如去冰箱。我们的方法不是获取环境的度量语义地图并使用导航规划,而是在捕获空间布局和语义上下文线索的表示之上学习导航策略。我们建议使用语义分割和检测掩码作为最先进的计算机视觉算法获得的观察结果,并使用深度网络来学习导航策略。模拟环境中公平表示的可用性使得使用真实和模拟数据进行联合训练成为可能,并减轻了通常用于解决学习策略从模拟到真实迁移的领域适应或领域随机化的需要。正如主动视觉数据集[1]所示,这种表示和导航策略都可以很容易地应用于真实的非合成环境。在未开发的环境中,我们的方法在54%的情况下成功达到目标,而非基于学习的方法为46%,基于学习的基线为28%。
What is a good visual representation for navigation? We study this question in the context of semantic visual navigation, which is the problem of a robot finding its way through a previously unseen environment to a target object, e.g. go to the refrigerator. Instead of acquiring a metric semantic map of an environment and using planning for navigation, our approach learns navigation policies on top of representations that capture spatial layout and semantic contextual cues. We propose to use semantic segmentation and detection masks as observations obtained by state-of-the-art computer vision algorithms and use a deep network to learn the navigation policy. The availability of equitable representations in simulated environments enables joint training using real and simulated data and alleviates the need for domain adaptation or domain randomization commonly used to tackle the sim-to-real transfer of the learned policies. Both the representation and the navigation policy can be readily applied to real non-synthetic environments as demonstrated on the Active Vision Dataset [1]. Our approach successfully gets to the target in 54% of the cases in unexplored environments, compared to 46% for a non-learning based approach, and 28% for a learning-based baseline.