Sim-to-Lab-to-Real: Safe Reinforcement Learning with Shielding and Generalization Guarantees

Sim-to-Lab-to-Real: Safe Reinforcement Learning with Shielding and Generalization Guarantees
复制标题

DOI:
10.1016/j.artint.2022.103811
复制
发表时间:
2022-01
期刊:
--
影响因子:
--
通讯作者:
Kai Hsu;Allen Z. Ren;D. Nguyen;Anirudha Majumdar;J. Fisac
Kai Hsu;Allen Z. Ren;D. Nguyen;Anirudha Majumdar;J. Fisac
中科院分区:
其他
文献类型:
--
作者:
Kai Hsu;Allen Z. Ren;D. Nguyen;Anirudha Majumdar;J. Fisac

文献摘要

相似文献

安全性是自主系统的关键组成部分,并且对于在真实的世界中使用基于学习的策略仍然是一个挑战。特别是,由于不安全的行为,使用强化学习学习的策略通常无法推广到新的环境。在本文中,我们提出了模拟到实验室到真实的弥合现实差距的概率保证安全意识的政策分布。为了提高安全性,我们应用了双重策略设置,其中使用累积任务奖励训练性能策略,并通过基于Hamilton Jacobi(HJ)可达性分析求解安全Bellman方程来训练备份(安全)策略。在Sim-to-Labtransfer中,我们应用监督控制方案来屏蔽探索过程中的不安全行为;在Lab-to-Realtransfer中,我们利用可能近似正确(PAC)-贝叶斯框架来提供未知环境中政策的预期性能和安全性的下限。此外,继承HJ可达性分析,该界限考虑了每个环境中最坏情况安全性的预期。我们实证研究的自我视觉导航在两种类型的室内环境中具有不同程度的照相现实主义的建议框架。我们还展示了强大的泛化性能,通过硬件实验在真实的室内空间与四足机器人。有关补充材料,请参见https://sites.google.com/princeton.edu/sim-to-lab-to-real。
Safety is a critical component of autonomous systems and remains a challenge for learning-based policies to be utilized in the real world. In particular, policies learned using reinforcement learning often fail to generalize to novel environments due to unsafe behavior. In this paper, we propose Sim-to-Lab-to-Real to bridge the reality gap with a probabilistically guaranteed safety-aware policy distribution. To improve safety, we apply a dual policy setup where a performance policy is trained using the cumulative task reward and a backup (safety) policy is trained by solving the Safety Bellman Equation based on Hamilton-Jacobi (HJ) reachability analysis. InSim-to-Labtransfer, we apply a supervisory control scheme to shield unsafe actions during exploration; inLab-to-Realtransfer, we leverage the Probably Approximately Correct (PAC)-Bayes framework to provide lower bounds on the expected performance and safety of policies in unseen environments. Additionally, inheriting from the HJ reachability analysis, the bound accounts for the expectation over the worst-case safety in each environment. We empirically study the proposed framework for ego-vision navigation in two types of indoor environments with varying degrees of photorealism. We also demonstrate strong generalization performance through hardware experiments in real indoor spaces with a quadrupedal robot. See https://sites.google.com/princeton.edu/sim-to-lab-to-real for supplementary material.