Safe Reinforcement Learning with Stability & Safety Guarantees Using Robust MPC

Safe Reinforcement Learning with Stability & Safety Guarantees Using Robust MPC
复制标题

安全稳定的强化学习

DOI:
--
复制
发表时间:
2020
期刊:
arXiv.org
影响因子:
--
通讯作者:
Mario Zanon
Mario Zanon
中科院分区:
--
文献类型:
--
作者:
S. Gros;Mario Zanon

文献摘要

被引文献

相似文献

强化学习提供了基于从服从策略的真实的系统获得的数据来优化策略的工具。虽然强化学习的潜力已经被很好地理解,但仍有许多关键方面需要解决。一个关键方面是安全和稳定问题。最近的出版物建议使用非线性模型预测控制技术结合强化学习作为解决这些问题的可行且理论上合理的方法。特别是,有人建议,鲁棒MPC允许在强化学习的背景下做出正式的稳定性和安全性声明。然而,仍然缺乏一个正式的理论,详细说明如何通过强化学习工具提供的参数更新来加强安全性和稳定性。本文将讨论这一差距。该理论是为通用的鲁棒MPC的情况下,并进一步详细的鲁棒管为基础的线性MPC的情况下,该理论是相当容易部署在实践中。
Reinforcement Learning offers tools to optimize policies based on the data obtained from the real system subject to the policy. While the potential of Reinforcement Learning is well understood, many critical aspects still need to be tackled. One crucial aspect is the issue of safety and stability. Recent publications suggest the use of Nonlinear Model Predictive Control techniques in combination with Reinforcement Learning as a viable and theoretically justified approach to tackle these problems. In particular, it has been suggested that robust MPC allows for making formal stability and safety claims in the context of Reinforcement Learning. However, a formal theory detailing how safety and stability can be enforced through the parameter updates delivered by the Reinforcement Learning tools is still lacking. This paper addresses this gap. The theory is developed for the generic robust MPC case, and further detailed in the robust tube-based linear MPC case, where the theory is fairly easy to deploy in practice.