Reachability-Based Trajectory Safeguard (RTS): A Safe and Fast Reinforcement Learning Safety Layer for Continuous Control

Reachability-Based Trajectory Safeguard (RTS): A Safe and Fast Reinforcement Learning Safety Layer for Continuous Control
复制标题

DOI:
10.1109/lra.2021.3063989
复制
发表时间:
2020-11
影响因子:
5.2
通讯作者:
Y. Shao;Chao Chen;Shreyas Kousik;Ram Vasudevan
Y. Shao;Chao Chen;Shreyas Kousik;Ram Vasudevan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Y. Shao;Chao Chen;Shreyas Kousik;Ram Vasudevan

文献摘要

被引文献

相似文献

强化学习(RL)算法通过试错法对长期累积奖励进行推理,在决策和控制任务中取得了显著的性能。然而,在强化学习训练期间,将这种试错方法应用于在安全关键环境中运行的真实世界机器人可能会导致碰撞。为了应对这一挑战,本文提出了一种基于可达性的轨迹保护(RTS)方法,该方法利用可达性分析来确保训练和操作过程中的安全性。给定一个已知(但不确定)的机器人模型,RTS预先计算机器人跟踪一系列参数化轨迹的正向可达集。在运行时,强化学习智能体以滚动时域的方式从该系列中进行选择以控制机器人;正向可达集用于识别智能体的选择是否安全,并对不安全的选择进行调整。通过在三个非线性机器人模型(包括一个12维四旋翼无人机)的静态环境中进行仿真,并与最先进的安全运动规划方法进行比较,说明了该方法的有效性。
Reinforcement Learning (RL) algorithms have achieved remarkable performance in decision making and control tasks by reasoning about long-term, cumulative reward using trial and error. However, during RL training, applying this trial-and-error approach to real-world robots operating in safety critical environment may lead to collisions. To address this challenge, this letter proposes a Reachability-based Trajectory Safeguard (RTS), which leverages reachability analysis to ensure safety during training and operation. Given a known (but uncertain) model of a robot, RTS precomputes a Forward Reachable Set of the robot tracking a continuum of parameterized trajectories. At runtime, the RL agent selects from this continuum in a receding-horizon way to control the robot; the FRS is used to identify if the agent's choice is safe or not, and to adjust unsafe choices. The efficacy of this method is illustrated in static environments on three nonlinear robot models, including a 12-D quadrotor drone, in simulation and in comparison with state-of-the-art safe motion planning methods.