Cooperative Finitely Excited Learning for Dynamical Games

Cooperative Finitely Excited Learning for Dynamical Games
复制标题

DOI:
10.1109/tcyb.2023.3274908
复制
发表时间:
2023-05
影响因子:
11.8
通讯作者:
Yongliang Yang;H. Modares;K. Vamvoudakis;F. Lewis
Yongliang Yang;H. Modares;K. Vamvoudakis;F. Lewis
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yongliang Yang;H. Modares;K. Vamvoudakis;F. Lewis

文献摘要

被引文献

相似文献

在本文中,我们提出了一种方法来增强零和博弈的学习框架,这种零和博弈具有连续时间的动态演化。与传统的集中式参与者-评论家学习方法不同,本文提出了一种新的合作有限激励学习方法,将在线记录数据与即时数据相结合,提高了学习效率。通过对每个智能体使用经验重放技术和智能体之间的分布式交互,我们能够用易于检查的合作激励条件取代经典的持续激励条件。该方法还保证了分布式参与者-评论家学习在Hamilton-Jacobi-Isaacs (HJI)方程解上的一致性。结果表明,该方法既能保证平衡点的闭环稳定性,又能保证收敛到纳什平衡点。仿真结果证明了该方法的有效性。
In this article, we propose a way to enhance the learning framework for zero-sum games with dynamics evolving in continuous time. In contrast to the conventional centralized actor–critic learning, a novel cooperative finitely excited learning approach is developed to combine the online recorded data with instantaneous data for efficiency. By using an experience replay technique for each agent and distributed interaction amongst agents, we are able to replace the classical persistent excitation condition with an easy-to-check cooperative excitation condition. This approach also guarantees the consensus of the distributed actor–critic learning on the solution to the Hamilton–Jacobi–Isaacs (HJI) equation. It is shown that both the closed-loop stability of the equilibrium point and convergence to the Nash equilibrium can be guaranteed. Simulation results demonstrate the efficacy of this approach compared to previous methods.