Safe and Efficient Reinforcement Learning using Disturbance-Observer-Based Control Barrier Functions

Safe and Efficient Reinforcement Learning using Disturbance-Observer-Based Control Barrier Functions
复制标题

DOI:
--
复制
发表时间:
2022-11
期刊:
--
影响因子:
--
通讯作者:
Yikun Cheng;Pan Zhao;N. Hovakimyan
Yikun Cheng;Pan Zhao;N. Hovakimyan
中科院分区:
其他
文献类型:
--
作者:
Yikun Cheng;Pan Zhao;N. Hovakimyan

文献摘要

相似文献

在训练过程中保证硬状态约束满足的安全强化学习(RL)近年来受到了广泛关注。安全过滤器,例如,基于控制屏障函数(CBFs),通过动态修改RL代理的不安全行为,为安全RL提供了一种有希望的方法。现有的基于安全滤波器的方法通常涉及对不确定动力学的学习和对学习到的模型误差的量化,这导致在收集大量数据以学习一个好的模型之前使用保守滤波器,从而阻碍了有效的探索。本文提出了一种利用扰动观测器(dob)和控制屏障函数(CBFs)实现安全高效RL的方法。与大多数现有的处理硬状态约束的安全强化学习方法不同,我们的方法不涉及模型学习,并利用dob准确地估计不确定性的点向值,然后将其合并到鲁棒CBF条件中以生成安全动作。基于dob的CBF可以作为无模型RL算法的安全过滤器,在必要时最小限度地修改RL代理的行为,以确保整个学习过程的安全性。在独轮车和2D四旋翼飞行器上的仿真结果表明,该方法在安全违规率、样本和计算效率方面优于使用CBFs和基于高斯过程的模型学习的最先进的安全RL算法。
Safe reinforcement learning (RL) with assured satisfaction of hard state constraints during training has recently received a lot of attention. Safety filters, e.g., based on control barrier functions (CBFs), provide a promising way for safe RL via modifying the unsafe actions of an RL agent on the fly. Existing safety filter-based approaches typically involve learning of uncertain dynamics and quantifying the learned model error, which leads to conservative filters before a large amount of data is collected to learn a good model, thereby preventing efficient exploration. This paper presents a method for safe and efficient RL using disturbance observers (DOBs) and control barrier functions (CBFs). Unlike most existing safe RL methods that deal with hard state constraints, our method does not involve model learning, and leverages DOBs to accurately estimate the pointwise value of the uncertainty, which is then incorporated into a robust CBF condition to generate safe actions. The DOB-based CBF can be used as a safety filter with model-free RL algorithms by minimally modifying the actions of an RL agent whenever necessary to ensure safety throughout the learning process. Simulation results on a unicycle and a 2D quadrotor demonstrate that the proposed method outperforms a state-of-the-art safe RL algorithm using CBFs and Gaussian processes-based model learning, in terms of safety violation rate, and sample and computational efficiency.