Socially aware robot navigation in crowds via deep reinforcement learning with resilient reward functions

Socially aware robot navigation in crowds via deep reinforcement learning with resilient reward functions
复制标题

DOI:
10.1080/01691864.2022.2043184
复制
发表时间:
2022-03
期刊:
影响因子:
2
通讯作者:
Xiaojun Lu;Hanwool Woo;Angela Faragasso;A. Yamashita;H. Asama
Xiaojun Lu;Hanwool Woo;Angela Faragasso;A. Yamashita;H. Asama
中科院分区:
计算机科学4区
文献类型:
--
作者:
Xiaojun Lu;Hanwool Woo;Angela Faragasso;A. Yamashita;H. Asama

文献摘要

相似文献

在机器人与人类共存的环境中导航的机器人不仅需要优化其路径,以实现与任务相关的性能(例如安全性和效率),而且还需要优化其对其他行人的社交顺应性。这是一项至关重要但具有挑战性的任务。以前的工作已经显示了深度强化学习(DRL)技术的力量,它通过使用它们来训练机器人导航的有效策略。然而,随着人群规模的增加,他们的表现会恶化。我们通过允许机器人与行人保持一段自适应距离来解决这个问题,并即使在高密度环境中也能执行安全导航。我们首先从一个真实的跟踪数据集中推导出一个表示不舒适距离与行人密度之间关系的定量公式。然后将该公式应用于DRL的报酬成形,得到具有弹性的报酬函数(R2F)。定性和定量的评估结果表明,在低和高行人密度环境下,我们的方法都优于最先进的方法。图形摘要
Robots navigating in a robot–human coexisting environment need to optimize their paths not only for task-related performance (e.g. safety and efficiency) but also for their social compliance to other pedestrians. This is a crucial yet challenging task. Previous work has shown the power of deep reinforcement learning (DRL) techniques by employing them to train efficient policies for robot navigation. However, their performance deteriorates when the crowd size grows. We cope with this problem by allowing the robot to keep an adapting distance from the pedestrians and perform safe navigation even in high density environments. We first derive a quantitative formula representing the relationship between uncomfortable distance and pedestrian density from a real-word tracking dataset. Then this formula is applied in reward shaping of DRL to get resilient reward functions (R2F). Qualitative and quantitative evaluation results demonstrate that our method outperforms state-of-the-art methods in both low and high pedestrian density environments. GRAPHICAL ABSTRACT