Can a Robot Trust You? : A DRL-Based Approach to Trust-Driven Human-Guided Navigation

Can a Robot Trust You? : A DRL-Based Approach to Trust-Driven Human-Guided Navigation
复制标题

机器人可以信任你吗?

DOI:
--
复制
发表时间:
2020
期刊:
IEEE International Conference on Robotics and Automation
影响因子:
--
通讯作者:
Aniket Bera
Aniket Bera
中科院分区:
--
文献类型:
--
作者:
Vishnu Sashank Dorbala;Arjun Srinivasan;Aniket Bera

文献摘要

参考文献

被引文献

相似文献

众所周知,人类利用各种感知输入构建日常环境的认知地图。因此,当人类被询问前往特定位置的方向时,他们将认知地图转换为方向指令的寻路能力受到挑战。由于空间焦虑,口头指示中使用的语言可能含糊且常常不清楚。为了解决导航指导中的这种不可靠性问题,我们提出了一种新颖的基于深度强化学习(DRL)的信任驱动机器人导航算法,该算法可以学习人类执行语言引导导航任务的可信度。我们的方法旨在回答机器人是否可以信任人类导航指导的问题。为此,我们着眼于训练一种策略,该策略学习仅使用值得信赖的人类指导来导航到目标位置,并由其自己的机器人信任指标驱动。我们着眼于量化基于语言的指令中的各种情感特征,并以人类信任指标的形式将它们纳入我们政策的观察空间中。我们将这两个信任指标运用到最佳认知推理方案中,该方案决定何时以及何时不信任给定的指导。我们的结果表明,学习的策略可以以最佳、省时的方式驾驭环境,而不是执行相同任务的探索性方法。我们展示了我们的结果在模拟和现实环境中的有效性。
Humans are known to construct cognitive maps of their everyday surroundings using a variety of perceptual inputs. As such, when a human is asked for directions to a particular location, their wayfinding capability in converting this cognitive map into directional instructions is challenged. Owing to spatial anxiety, the language used in the spoken instructions can be vague and often unclear. To account for this unreliability in navigational guidance, we propose a novel Deep Reinforcement Learning (DRL) based trust-driven robot navigation algorithm that learns humans’ trustworthiness to perform a language guided navigation task.Our approach seeks to answer the question as to whether a robot can trust a human’s navigational guidance or not. To this end, we look at training a policy that learns to navigate towards a goal location using only trustworthy human guidance, driven by its own robot trust metric. We look at quantifying various affective features from language-based instructions and incorporate them into our policy’s observation space in the form of a human trust metric. We utilize both these trust metrics into an optimal cognitive reasoning scheme that decides when and when not to trust the given guidance. Our results show that the learned policy can navigate the environment in an optimal, time-efficient manner as opposed to an explorative approach that performs the same task. We showcase the efficacy of our results both in simulation and a real world environment.
DOI: 10.15607/rss.2020.xvi.102
发表时间: 2020-07
期刊: Robotics: Science and Systems XVI
影响因子: --
作者:
N. Gopalan;Eric Rosen;G. Konidaris;Stefanie Tellex
通讯作者: N. Gopalan;Eric Rosen;G. Konidaris;Stefanie Tellex