Online Reinforcement Learning Control by Direct Heuristic Dynamic Programming: From Time-Driven to Event-Driven

Online Reinforcement Learning Control by Direct Heuristic Dynamic Programming: From Time-Driven to Event-Driven
复制标题

DOI:
10.1109/tnnls.2021.3053037
复制
发表时间:
2020-06
影响因子:
10.4
通讯作者:
Qingtao Zhao;J. Si;Jian Sun-
Qingtao Zhao;J. Si;Jian Sun-
中科院分区:
计算机科学1区
文献类型:
--
作者:
Qingtao Zhao;J. Si;Jian Sun-

文献摘要

被引文献

相似文献

在这项工作中,时间驱动学习是指随着新数据的到来而不断更新预测模型中的参数的机器学习方法。在现有的近似动态规划(ADP)和强化学习(RL)算法中,直接启发式动态规划(dHDP)已被证明是解决几个复杂学习控制问题的有效工具。随着系统状态的不断发展,它不断更新控制策略和批评家。因此,需要防止时间驱动的 dHDP 由于诸如噪声之类的无关紧要的系统事件而更新。为了实现这一目标,我们提出了一种新的事件驱动的 dHDP。通过构造李亚普诺夫候选函数,我们证明了系统状态的一致最终有界性(UUB)以及批评家和控制策略网络中的权重。因此,我们展示了在有限范围内接近贝尔曼最优性的近似控制和成本函数。我们还说明了事件驱动 dHDP 算法与原始时间驱动 dHDP 算法的工作原理。
In this work, time-driven learning refers to the machine learning method that updates parameters in a prediction model continuously as new data arrives. Among existing approximate dynamic programming (ADP) and reinforcement learning (RL) algorithms, the direct heuristic dynamic programming (dHDP) has been shown an effective tool as demonstrated in solving several complex learning control problems. It continuously updates the control policy and the critic as system states continuously evolve. It is therefore desirable to prevent the time-driven dHDP from updating due to insignificant system event such as noise. Toward this goal, we propose a new event-driven dHDP. By constructing a Lyapunov function candidate, we prove the uniformly ultimately boundedness (UUB) of the system states and the weights in the critic and the control policy networks. Consequently, we show the approximate control and cost-to-go function approaching Bellman optimality within a finite bound. We also illustrate how the event-driven dHDP algorithm works in comparison to the original time-driven dHDP.