Feedback-efficient Active Preference Learning for Socially Aware Robot Navigation

Feedback-efficient Active Preference Learning for Socially Aware Robot Navigation
复制标题

DOI:
10.1109/iros47612.2022.9981616
复制
发表时间:
2022-01
期刊:
2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Ruiqi Wang;Weizheng Wang;Byung-Cheol Min
Ruiqi Wang;Weizheng Wang;Byung-Cheol Min
中科院分区:
其他
文献类型:
--
作者:
Ruiqi Wang;Weizheng Wang;Byung-Cheol Min

文献摘要

被引文献

相似文献

具有社会意识的机器人导航,其中机器人需要优化其轨迹,以保持舒适和兼容的空间与人类的互动,除了达到其目标,没有碰撞,是一个基本的,但具有挑战性的任务,在人机交互的背景下。虽然现有的基于学习的方法比以前的基于模型的方法取得了更好的性能,但它们仍然存在缺点:强化学习依赖于手工制作的奖励,不太可能有效地量化广泛的社会遵从性,并可能导致奖励剥削问题;同时,逆向强化学习需要昂贵的人类演示。在本文中,我们提出了一个反馈有效的主动偏好学习方法,FAPL,蒸馏人类的舒适度和期望到一个奖励模型,以指导机器人代理探索潜在的社会合规方面。我们进一步引入混合经验学习,以提高人类的反馈和样本的效率,并通过广泛的仿真实验和用户研究(N=10)采用物理机器人与人类主体在现实世界的场景中导航从FAPL学习的机器人行为的好处进行评估。有关这项工作的源代码和实验视频,请访问:https://sites.google.com/view/san-fapl。
Socially aware robot navigation, where a robot is required to optimize its trajectory to maintain comfortable and compliant spatial interactions with humans in addition to reaching its goal without collisions, is a fundamental yet challenging task in the context of human-robot interaction. While existing learning-based methods have achieved better performance than the preceding model-based ones, they still have drawbacks: reinforcement learning depends on the handcrafted reward that is unlikely to effectively quantify broad social compliance, and can lead to reward exploitation problems; meanwhile, inverse rein-forcement learning suffers from the need for expensive human demonstrations. In this paper, we propose a feedback-efficient active preference learning approach, FAPL, that distills human comfort and expectation into a reward model to guide the robot agent to explore latent aspects of social compliance. We further introduce hybrid experience learning to improve the efficiency of human feedback and samples, and evaluate benefits of robot behaviors learned from FAPL through extensive simulation experiments and a user study (N=10) employing a physical robot to navigate with human subjects in real-world scenarios. Source code and experiment videos for this work are available at: https://sites.google.com/view/san-fapl.