Using process data to generate an optimal control policy via apprenticeship and reinforcement learning

Using process data to generate an optimal control policy via apprenticeship and reinforcement learning
复制标题

使用过程数据通过学徒和强化学习生成最优控制策略

DOI:
10.1002/aic.17306
复制
发表时间:
2021-05-15
期刊:
影响因子:
3.7
通讯作者:
Zhang, Dongda
Zhang, Dongda
中科院分区:
工程技术3区
文献类型:
--
作者:
Mowbray, Max;Smith, Robin;Zhang, Dongda

文献摘要

被引文献

相似文献

强化学习(RL)是一种数据驱动的综合最优控制策略的方法。基于强化学习的控制器广泛应用的一个障碍是它在在线训练期间对数据的渴求,以及它无法从人工操作员和历史过程操作数据中提取有用的信息。在这里,我们提出了一个两步框架来解决这一挑战。首先,我们通过逆强化学习采用学徒学习来分析历史过程数据,以同步识别奖励函数和参数化控制策略。这是离线进行的。其次,在持续过程中,通过强化学习,只需几次迭代,即可在线有效地改进参数化。该框架的显著优点包括允许热启动RL算法进行过程最优控制,以及对现有控制器和数据控制知识的鲁棒抽象。该框架通过三个案例研究进行了演示,显示了其在化学过程控制方面的潜力。
Reinforcement learning (RL) is a data-driven approach to synthesizing an optimal control policy. A barrier to wide implementation of RL-based controllers is its data-hungry nature during online training and its inability to extract useful information from human operator and historical process operation data. Here, we present a two-step framework to resolve this challenge. First, we employ apprenticeship learning via inverse RL to analyze historical process data for synchronous identification of a reward function and parameterization of the control policy. This is conducted offline. Second, the parameterization is improved online efficiently under the ongoing process via RL within only a few iterations. Significant advantages of this framework include to allow for the hot-start of RL algorithms for process optimal control, and robust abstraction of existing controllers and control knowledge from data. The framework is demonstrated on three case studies, showing its potential for chemical process control.