Bridging Model-based Safety and Model-free Reinforcement Learning through System Identification of Low Dimensional Linear Models

Bridging Model-based Safety and Model-free Reinforcement Learning through System Identification of Low Dimensional Linear Models
复制标题

DOI:
10.48550/arxiv.2205.05787
复制
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhongyu Li;Jun Zeng;A. Thirugnanam;K. Sreenath
Zhongyu Li;Jun Zeng;A. Thirugnanam;K. Sreenath
中科院分区:
其他
文献类型:
--
作者:
Zhongyu Li;Jun Zeng;A. Thirugnanam;K. Sreenath

文献摘要

相似文献

基于模型的方法能够提供形式化的安全保证,而基于模型的方法能够通过学习全阶系统动力学来开发机器人的敏捷性,因此将基于模型的安全性和无模型强化学习(RL)相结合的方法在动态机器人中引起了广泛的关注。然而,目前解决这个问题的方法大多局限于简单的系统。在本文中,我们提出了一种将基于模型的安全性和无模型强化学习相结合的新方法,方法是显式地找到受RL策略控制的系统的低维模型,并对该简单模型进行稳定性和安全性保证。我们以复杂的两足机器人CASSIE及其基于RL的步行控制器为例,CASSIE是一个具有混合动力学和欠驱动的高维非线性系统。我们证明一个低维的动力学模型足以捕捉闭环系统的动力学。我们证明了该模型是线性的,渐近稳定的,并且在所有维度的控制输入上是解耦的。我们进一步举例说明,即使在使用不同的RL控制策略时,也存在这样的线性。这些结果为理解RL和最优控制之间的关系指出了一个有趣的方向:在某些情况下,RL是否倾向于在训练过程中线性化非线性系统。此外,以CASSIE自主导航为例,利用基于RL的控制器所提供的敏捷性,证明了所建立的线性模型能够通过安全关键的最优控制框架,例如具有控制屏障功能的模型预测控制来提供保证。
Bridging model-based safety and model-free reinforcement learning (RL) for dynamic robots is appealing since model-based methods are able to provide formal safety guarantees, while RL-based methods are able to exploit the robot agility by learning from the full-order system dynamics. However, current approaches to tackle this problem are mostly restricted to simple systems. In this paper, we propose a new method to combine model-based safety with model-free reinforcement learning by explicitly finding a low-dimensional model of the system controlled by a RL policy and applying stability and safety guarantees on that simple model. We use a complex bipedal robot Cassie, which is a high dimensional nonlinear system with hybrid dynamics and underactuation, and its RL-based walking controller as an example. We show that a low-dimensional dynamical model is sufficient to capture the dynamics of the closed-loop system. We demonstrate that this model is linear, asymptotically stable, and is decoupled across control input in all dimensions. We further exemplify that such linearity exists even when using different RL control policies. Such results point out an interesting direction to understand the relationship between RL and optimal control: whether RL tends to linearize the nonlinear system during training in some cases. Furthermore, we illustrate that the found linear model is able to provide guarantees by safety-critical optimal control framework, e.g., Model Predictive Control with Control Barrier Functions, on an example of autonomous navigation using Cassie while taking advantage of the agility provided by the RL-based controller.