Learning Adaptive Optimal Controllers for Linear Time-Delay Systems *

Learning Adaptive Optimal Controllers for Linear Time-Delay Systems *
复制标题

DOI:
10.23919/acc55779.2023.10156108
复制
发表时间:
2023-05
期刊:
2023 American Control Conference (ACC)
影响因子:
--
通讯作者:
Leilei Cui;Bo Pang;Zhong-Ping Jiang
Leilei Cui;Bo Pang;Zhong-Ping Jiang
中科院分区:
其他
文献类型:
--
作者:
Leilei Cui;Bo Pang;Zhong-Ping Jiang

文献摘要

相似文献

研究了一类无穷维线性时滞系统的基于学习的最优控制问题。目的是填补自适应动态规划(ADP)的差距,其中自适应最优控制的无限维系统没有解决。将经典的基于模型的时滞系统线性二次型(LQ)最优控制与最新的强化学习(RL)技术相结合是联合收割机的关键策略。基于模型和数据驱动的策略迭代(PI)的方法被提出来解决相应的代数Riccati方程(ARE)的保证收敛。所提出的PI算法可以被认为是ADP到无穷维时滞系统的推广。混合交通环境下的自动驾驶中,考虑人类驾驶员的反应延迟所产生的实际应用表明,所提出的算法的效率。
This paper studies the learning-based optimal control for a class of infinite-dimensional linear time-delay systems. The aim is to fill the gap of adaptive dynamic programming (ADP) where adaptive optimal control of infinite-dimensional systems is not addressed. A key strategy is to combine the classical model-based linear quadratic (LQ) optimal control of time-delay systems with the state-of-art reinforcement learning (RL) technique. Both the model-based and data-driven policy iteration (PI) approaches are proposed to solve the corresponding algebraic Riccati equation (ARE) with guaranteed convergence. The proposed PI algorithm can be considered as a generalization of ADP to infinite-dimensional time-delay systems. The efficiency of the proposed algorithm is demonstrated by the practical application arising from autonomous driving in mixed traffic environments, where human drivers’ reaction delay is considered.