Separation of learning and control for cyber-physical systems

Separation of learning and control for cyber-physical systems
复制标题

DOI:
10.1016/j.automatica.2023.110912
复制
发表时间:
2021-07
期刊:
Autom.
影响因子:
--
通讯作者:
Andreas A. Malikopoulos
Andreas A. Malikopoulos
中科院分区:
其他
文献类型:
--
作者:
Andreas A. Malikopoulos

文献摘要

被引文献

相似文献

大多数网络物理系统(CP)会遇到大量的数据,这些数据是实时地逐渐添加到系统中的,而不是提前完全添加到系统中的。在本文中,我们提供了一个理论框架,在控制理论和学习的交叉点上,给出了这类CP的最优控制策略。在提议的框架中,我们使用实际的CPS,即我们寻求在线最优控制的“真实”系统,与可用的CPS模型并行。然后,我们为系统建立一个不依赖于控制策略的信息状态。这种独立性的一个重要结果是,对于任何给定的控制策略的选择和直到时间t的系统变量的实现,未来时间的信息状态不取决于在时间t的控制策略的选择,而只取决于在时间t的决策的实现,因此它们与状态估计和控制之间的分离的概念有关。也就是说,未来信息状态与当前控制策略的选择是分离的。这种控制策略称为分离控制策略。因此,我们可以离线推导出系统关于信息状态的最优控制策略,然后使用标准学习方法在线学习信息状态,同时数据逐渐实时地添加到系统中。证明了在已知信息状态后,离线推导出的CPS模型的分离控制策略对实际系统是最优的。我们在一个由两个具有延迟共享信息结构的子系统组成的动态系统中说明了所提出的框架。
Most cyber–physical systems (CPS) encounter a large volume of data which is added to the system gradually in real time and not altogether in advance. In this paper, we provide a theoretical framework that yields optimal control strategies for such CPS at the intersection of control theory and learning. In the proposed framework, we use the actual CPS, ie, the “true” system that we seek to optimally control online, in parallel with a model of the CPS that is available. We then institute an information state for the system which does not depend on the control strategy. An important consequence of this independence is that for any given choice of a control strategy and a realization of the system’s variables until time t, the information states at future times do not depend on the choice of the control strategy at time t but only on the realization of the decision at time t, and thus they are related to the concept of separation between estimation of the state and control. Namely, the future information states are separated from the choice of the current control strategy. Such control strategies are called separated control strategies. Hence, we can derive offline the optimal control strategy of the system with respect to the information state, which might not be precisely known due to model uncertainties or complexity of the system, and then use standard learning approaches to learn the information state online while data are added gradually to the system in real time. We show that after the information state becomes known, the separated control strategy of the CPS model derived offline is optimal for the actual system. We illustrate the proposed framework in a dynamic system consisting of two subsystems with a delayed sharing information structure.