Data-Driven Economic NMPC Using Reinforcement Learning

Data-Driven Economic NMPC Using Reinforcement Learning
复制标题

DOI:
10.1109/tac.2019.2913768
复制
发表时间:
2019-04
影响因子:
6.8
通讯作者:
S. Gros;Mario Zanon
S. Gros;Mario Zanon
中科院分区:
计算机科学2区
文献类型:
--
作者:
S. Gros;Mario Zanon

文献摘要

被引文献

相似文献

强化学习(RL)是一种强大的工具,可以在不依赖系统模型的情况下执行数据驱动的最优控制。然而,RL努力为最终控制方案的行为提供硬保证。相比之下,非线性模型预测控制(NMPC)和经济NMPC(ENMPC)是具有约束和限制的复杂系统的闭环最优控制的标准工具,并受益于丰富的理论来评估其闭环行为。不幸的是,(E)NMPC的性能取决于控制方案所基于的模型的质量。在本文中,我们表明,(E)NMPC计划可以调整提供真实的系统的最优策略,即使在使用错误的模型。这一结果也适用于具有随机动态的真实的系统。这使得ENMPC可以用作RL中的新型函数逼近器。此外,我们调查我们的结果在ENMPC的上下文中,并正式连接到耗散性的概念,这是中央的ENMPC稳定性。最后,我们详细介绍了如何使用这些结果来部署经典的RL工具来调整(E)NMPC方案。我们应用这些工具,一个经典的线性MPC设置和一个标准的非线性例子,从ENMPC文献。
Reinforcement learning (RL) is a powerful tool to perform data-driven optimal control without relying on a model of the system. However, RL struggles to provide hard guarantees on the behavior of the resulting control scheme. In contrast, nonlinear model predictive control (NMPC) and economic NMPC (ENMPC) are standard tools for the closed-loop optimal control of complex systems with constraints and limitations, and benefit from a rich theory to assess their closed-loop behavior. Unfortunately, the performance of (E)NMPC hinges on the quality of the model underlying the control scheme. In this paper, we show that an (E)NMPC scheme can be tuned to deliver the optimal policy of the real system even when using a wrong model. This result also holds for real systems having stochastic dynamics. This entails that ENMPC can be used as a new type of function approximator within RL. Furthermore, we investigate our results in the context of ENMPC and formally connect them to the concept of dissipativity, which is central for the ENMPC stability. Finally, we detail how these results can be used to deploy classic RL tools for tuning (E)NMPC schemes. We apply these tools on both, a classical linear MPC setting and a standard nonlinear example, from the ENMPC literature.