Model-Free Temporal Difference Learning for Non-Zero-Sum Games

Model-Free Temporal Difference Learning for Non-Zero-Sum Games
复制标题

DOI:
10.1109/ijcnn.2019.8851866
复制
发表时间:
2019-07
期刊:
2019 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Liming Wang;Yongliang Yang;Dawei Ding;Yixin Yin;Zhishan Guo;D. Wunsch
Liming Wang;Yongliang Yang;Dawei Ding;Yixin Yin;Zhishan Guo;D. Wunsch
中科院分区:
其他
文献类型:
--
作者:
Liming Wang;Yongliang Yang;Dawei Ding;Yixin Yin;Zhishan Guo;D. Wunsch

文献摘要

相似文献

研究了一类连续时间线性动态系统的二人非零和对策问题。证明了非零和博弈问题的结果是求解耦合代数Riccati方程,即非线性代数矩阵方程。与单参与者线性动力系统的代数Riccati方程相比,多参与者非零和博弈的耦合代数Riccati方程更难直接求解。首先,引入策略迭代算法求解非零和博弈的纳什均衡,这是求解耦合代数Riccati方程的充要条件。然而,策略迭代算法是离线的,需要完整的系统动力学知识。为了克服上述问题,提出了一种新的在线迭代算法——积分时间差分学习算法。此外,还给出了积分时间差分学习算法的等价紧凑形式。结果表明,积分时间差分学习算法可以在线实现,并且只需要了解系统动力学的部分知识。此外,在每个迭代步骤中,分析了使用积分时间差分学习算法的闭环稳定性。最后,通过仿真研究验证了该算法的有效性。
In this paper, we consider the two-player nonzero-sum games problem for continuous-time linear dynamic systems. It is shown that the non-zero-sum games problem results in solving the coupled algebraic Riccati equations, which are nonlinear algebraic matrix equations. Compared with the algebraic Riccati equation of the linear dynamic systems with only one player, the coupled algebraic Riccati equations of nonzero-sum games with multi-player are more difficult to be solved directly. First, the policy iteration algorithm is introduced to find the Nash equilibrium of the non-zero-sum games, which is the sufficient and necessary condition to solve the coupled algebraic Riccati equations. However, the policy iteration algorithm is offline and requires complete knowledge of the system dynamics. To overcome the above issues, a novel online iterative algorithm, named integral temporal difference learning algorithm, is developed. Moreover, an equivalent compact form of the integral temporal difference learning algorithm is also presented. It is shown that the integral temporal difference learning algorithm can be implemented in an online fashion and requires only partial knowledge of the system dynamics. In addition, in each iteration step, the closed-loop stability using the integral temporal difference learning algorithm is analyzed. Finally, the simulation study shows the effectiveness of the presented algorithm.