Probabilistic Safeguard for Reinforcement Learning Using Safety Index Guided Gaussian Process Models

Probabilistic Safeguard for Reinforcement Learning Using Safety Index Guided Gaussian Process Models
复制标题

DOI:
10.48550/arxiv.2210.01041
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Weiye Zhao;Tairan He;Changliu Liu
Weiye Zhao;Tairan He;Changliu Liu
中科院分区:
其他
文献类型:
--
作者:
Weiye Zhao;Tairan He;Changliu Liu

文献摘要

相似文献

安全是将强化学习 (RL) 应用到物理世界时最关心的问题之一。在其核心部分,确保 RL 代理在没有白盒或黑盒动力学模型的情况下持续满足硬状态约束是一项挑战。本文提出了一个集成的模型学习和安全控制框架来保护任何代理,其中其动态被学习为高斯过程。所提出的理论提供了(i)一种构建离线数据集的新方法,用于模型学习,最好地实现安全要求; (ii)安全指标参数化规则,保证安全控制的存在; (iii) 当使用上述数据集学习模型时,在概率前向不变性方面的安全保证。仿真结果表明,我们的框架保证了各种连续控制任务的安全违规几乎为零。
Safety is one of the biggest concerns to applying reinforcement learning (RL) to the physical world. In its core part, it is challenging to ensure RL agents persistently satisfy a hard state constraint without white-box or black-box dynamics models. This paper presents an integrated model learning and safe control framework to safeguard any agent, where its dynamics are learned as Gaussian processes. The proposed theory provides (i) a novel method to construct an offline dataset for model learning that best achieves safety requirements; (ii) a parameterization rule for safety index to ensure the existence of safe control; (iii) a safety guarantee in terms of probabilistic forward invariance when the model is learned using the aforementioned dataset. Simulation results show that our framework guarantees almost zero safety violation on various continuous control tasks.