Stochastic Two-Player Zero-Sum Learning Differential Games

Stochastic Two-Player Zero-Sum Learning Differential Games
复制标题

DOI:
10.1109/icca.2019.8899568
复制
发表时间:
2019-07
期刊:
2019 IEEE 15th International Conference on Control and Automation (ICCA)
影响因子:
--
通讯作者:
Mushuang Liu;Yan Wan;F. Lewis;V. Lopez
Mushuang Liu;Yan Wan;F. Lewis;V. Lopez
中科院分区:
其他
文献类型:
--
作者:
Mushuang Liu;Yan Wan;F. Lewis;V. Lopez

文献摘要

相似文献

两人零和微分对策已经被广泛研究,部分原因是它的解隐含了H_{\infty}$最优性。现有的零和微分对策的研究要么假设确定性的动态或被加性噪声破坏的动态。在现实环境中,高维环境的不确定性往往调制系统动力学在一个更复杂的方式。本文研究了由更一般的不确定线性动力系统支配的随机二人零和微分对策。我们表明,这个游戏的最优控制策略,可以通过求解Hamilton-Jacobi-Bellman(HJB)方程。我们证明了,与派生的最优控制策略,系统是渐近稳定的平均,并达到纳什均衡。针对随机二人零和博弈问题,设计了一种新的策略迭代(PI)算法,该算法结合了积分强化学习(IRL)和一种有效的不确定性评估方法--多元概率配置法(MPCM).该算法提供了一种快速的在线求解随机两人零和微分对策的系统动力学的多重不确定性。
The two-player zero-sum differential game has been extensively studied, partially because its solution implies the $H_{\infty}$ optimality. Existing studies on zero-sum differential games either assume deterministic dynamics or the dynamics corrupted by additive noise. In realistic environments, high-dimensional environmental uncertainties often modulate system dynamics in a more complicated fashion. In this paper, we study the stochastic two-player zero-sum differential game governed by more general uncertain linear dynamics. We show that the optimal control policies for this game can be found by solving the Hamilton-Jacobi-Bellman (HJB) equation. We prove that with the derived optimal control policies, the system is asymptotically stable in the mean, and reaches the Nash equilibrium. To solve the stochastic two-player zero-sum game online, we design a new policy iteration (PI) algorithm that integrates the integral reinforcement learning (IRL) and an efficient uncertainty evaluation method—multivariate probabilistic collocation method (MPCM). This algorithm provides a fast online solution for the stochastic two-player zero-sum differential game subject to multiple uncertainties in the system dynamics.