Enhancing Safety in Learning from Demonstration Algorithms via Control Barrier Function Shielding

Enhancing Safety in Learning from Demonstration Algorithms via Control Barrier Function Shielding
复制标题

DOI:
10.1145/3610977.3635002
复制
发表时间:
2024-03
期刊:
Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction
影响因子:
--
通讯作者:
Yue Yang;Letian Chen;Z. Zaidi;Sanne van Waveren;Arjun Krishna;M. Gombolay
Yue Yang;Letian Chen;Z. Zaidi;Sanne van Waveren;Arjun Krishna;M. Gombolay
中科院分区:
其他
文献类型:
--
作者:
Yue Yang;Letian Chen;Z. Zaidi;Sanne van Waveren;Arjun Krishna;M. Gombolay

文献摘要

相似文献

从演示中学习(LfD)是非机器人专家最终用户教授机器人新任务的一种强大方法,使他们能够定制机器人行为。然而,现代LfD技术没有明确地合成安全的机器人行为,这限制了这些方法在真实的世界中的可部署性。为了在不依赖专家的情况下加强LfD的安全性,我们提出了一个新的框架,SELDING与反向REINFORMATION学习(SECURE)中的控制障碍函数,它从最终用户那里学习定制的控制障碍函数(CBF),防止机器人采取不安全的行动,同时对任务完成施加很少的干扰。我们在三组实验中评估SECURE。首先,我们通过经验验证SECURE从演示中学习高质量的CBF,并在模拟机器人和自动驾驶任务中优于传统的LfD方法,安全性提高高达100%。其次,我们证明了机器人专家可以利用SECURE在现实世界的刀切、做饭任务中比传统的LfD方法完成12.5%的任务,同时将安全违规的数量降为零。最后,我们在一项用户研究中证明,非机器人专家可以使用SECURE来有效地教授机器人安全策略,避免与人发生碰撞,并防止咖啡溢出。
Learning from Demonstration (LfD) is a powerful method for non-roboticists end-users to teach robots new tasks, enabling them to customize the robot behavior. However, modern LfD techniques do not explicitly synthesize safe robot behavior, which limits the deployability of these approaches in the real world. To enforce safety in LfD without relying on experts, we propose a new framework, SElding with Control barrier fUnctions in inverse REinforcement learning (SECURE), which learns a customized Control Barrier Function (CBF) from end-users that prevents robots from taking unsafe actions while imposing little interference with the task completion. We evaluate SECURE in three sets of experiments. First, we empirically validate SECURE learns a high-quality CBF from demonstrations and outperforms conventional LfD methods on simulated robotic and autonomous driving tasks with improvements on safety by up to 100%. Second, we demonstrate that roboticists can leverage SECURE to outperform conventional LfD approaches on a real-world knife-cutting, meal-preparation task by 12.5% in task completion while driving the number of safety violations to zero. Finally, we demonstrate in a user study that non-roboticists can use SECURE to effectively teach the robot safe policies that avoid collisions with the person and prevent coffee from spilling.