Fast Real-Time Reinforcement Learning for Partially-Observable Large-Scale Systems

Fast Real-Time Reinforcement Learning for Partially-Observable Large-Scale Systems
复制标题

DOI:
10.1109/tai.2021.3058228
复制
发表时间:
2020-12
期刊:
IEEE Transactions on Artificial Intelligence
影响因子:
--
通讯作者:
T. Sadamoto;A. Chakrabortty
T. Sadamoto;A. Chakrabortty
中科院分区:
其他
文献类型:
--
作者:
T. Sadamoto;A. Chakrabortty

文献摘要

相似文献

针对大规模部分可观测线性动态系统,提出了一种快速实时强化学习控制算法。首先,我们开发了一个一次性的RL方法设计无模型的最优控制器的输入和输出的有限时间的历史的基础上。然而,当系统维数很大时,这种方法可能会遭受很长的学习时间。为了克服这个问题,在本文的后半部分,我们引入了一个新的概念,近似的设计,其中输入输出历史的原始集合被替换为一个更短的集合。我们表明,这种近似可以导致一个接近最优的控制器,是基于一个低维近似的原始系统的可达性和可观测性。我们提供了一个指导方针,以确定一个适当的长度的投入产出的历史,以减少次优差距。所得到的次优控制器的维数远小于最优控制器的维数,从而加快了学习时间。然而,学习的控制器,可能会导致不稳定时,在原来的高维系统中实施的近似误差的不利激励。我们从理论上建立了闭环稳定的条件,使用鲁棒控制理论,其次是学习时间,输入/输出历史的长度,和闭环性能之间的权衡数值研究。电力系统的部分可观测的非线性微分代数方程建模的例子说明了该方法的有效性。
We propose a fast real-time reinforcement learning (RL) control algorithm for large-scale partially-observable linear dynamic systems. We first develop a one-shot RL method for designing model-free optimal controllers based on a finite-time history of the inputs and the outputs. However, when the system dimension is large, this method may suffer from a long learning time. To overcome this problem, in the second half of the paper we introduce a new notion of approximation to the design, where the original set of input-output history is replaced by a much shorter set. We show that this approximation can lead to a nearly optimal controller that is based on a lower-dimensional approximant of the original system in terms of reachability and observability. We provide a guideline for determining an appropriate length of the input-output history to reduce the suboptimality gap. The dimension of the resulting suboptimal controller is far less than that of the optimal controller, thereby speeding up learning time. The learned controller, however may cause instability when implemented in the original high-dimensional system by adversely exciting the approximation error. We theoretically establish the conditions for closed-loop stability using robust control theory, followed by numerical investigations of the trade-offs between learning time, length of input/output history, and closed-loop performance. The effectiveness of the method is illustrated using examples from electric power systems, modeled by partially-observable nonlinear differential-algebraic equations.