A General Safety Framework for Learning-Based Control in Uncertain Robotic Systems

A General Safety Framework for Learning-Based Control in Uncertain Robotic Systems
复制标题

DOI:
10.1109/tac.2018.2876389
复制
发表时间:
2019-07-01
影响因子:
6.8
通讯作者:
Tomlin, Claire J.
Tomlin, Claire J.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Fisac, Jaime F.;Akametalu, Anayo K.;Tomlin, Claire J.

文献摘要

被引文献

相似文献

基于学习的控制方案已被证明是有效的,这有力地推动了它们在物理世界中操作的机器人系统中的应用。然而,在学习过程中保证正确的操作目前是一个悬而未决的问题,这是至关重要的安全关键系统。我们提出了一个通用的安全框架的基础上,汉密尔顿-雅可比可达性方法,可以与任意的学习算法。该方法利用系统动力学的近似知识,以保证约束满足,同时最小限度地干扰学习过程。我们进一步引入了贝叶斯机制,该机制可以在系统获得新证据时细化安全性分析,在适当的时候减少初始保守性,同时通过实时验证加强保证。其结果是一个最少的限制,安全保护的控制律,干预只有当计算的安全保证需要它,或在计算的保证的信心根据新的观察衰减。我们证明了理论上的安全保证相结合的概率和最坏情况下的分析,并证明了所提出的框架实验上的四旋翼飞行器。尽管安全性分析是基于简单的点质量模型,但四旋翼飞机通过策略梯度强化学习成功地达到了合适的控制器,而不会坠毁,并安全地从飞行过程中引入的强烈外部干扰中缩回。
The proven efficacy of learning-based control schemes strongly motivates their application to robotic systems operating in the physical world. However, guaranteeing correct operation during the learning process is currently an unresolved issue, which is of vital importance in safety-critical systems. We propose a general safety framework based on Hamilton-Jacobi reachability methods that can work in conjunction with an arbitrary learning algorithm. The method exploits approximate knowledge of the system dynamics to guarantee constraint satisfaction while minimally interfering with the learning process. We further introduce a Bayesian mechanism that refines the safety analysis as the system acquires new evidence, reducing initial conservativeness when appropriate while strengthening guarantees through real-time validation. The result is a least-restrictive, safety-preserving control law that intervenes only when the computed safety guarantees require it, or confidence in the computed guarantees decays in light of new observations. We prove theoretical safety guarantees combining probabilistic and worst-case analysis and demonstrate the proposed framework experimentally on a quadrotor vehicle. Even though safety analysis is based on a simple point-mass model, the quadrotor successfully arrives at a suitable controller by policy-gradient reinforcement learning without ever crashing, and safely retracts away from a strong external disturbance introduced during flight.