Integral Reinforcement Learning for online computation of feedback Nash strategies of nonzero-sum differential games

Integral Reinforcement Learning for online computation of feedback Nash strategies of nonzero-sum differential games
复制标题

DOI:
10.1109/cdc.2010.5718152
复制
发表时间:
2010-12
期刊:
49th IEEE Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
D. Vrabie;F. Lewis
D. Vrabie;F. Lewis
中科院分区:
其他
文献类型:
--
作者:
D. Vrabie;F. Lewis

文献摘要

被引文献

相似文献

本文提出了一种近似/自适应动态规划(ADP)算法,该算法在线求解具有线性动态和无限时域二次成本的两人非零和微分对策的纳什均衡.每个游戏参与者都使用积分强化学习(IRL)的过程来在线计算与每个给定的反馈控制策略集相关联的无限时域值函数。它将表明,在线算法在数学上是等效的离线迭代方法,以前在文献中介绍的,解决了一组耦合代数Riccati方程(ARE)的游戏问题,使用完整的知识系统动态。在这里,我们将展示如何ADP技术将提高离线方法的能力,允许在线解决方案,而不需要系统动态的完整知识。连续时间微分博弈中的两个参与者实时竞争,并根据系统的在线测量数据确定反馈纳什控制策略。该算法建立在学习阶段和策略更新步骤之间的相互作用上,在学习阶段,每个玩家在线学习他们与给定的一组游戏策略相关联的值,策略更新步骤由每个支付者执行,以降低他们的成本值。球员们同时学习。仿真结果证明了ADP方案的可行性。
This paper presents an Approximate/Adaptive Dynamic Programming (ADP) algorithm that finds online the Nash equilibrium for two-player nonzero-sum differential games with linear dynamics and infinite horizon quadratic cost. Each of the game players is using the procedure of Integral Reinforcement Learning (IRL) to calculate online the infinite horizon value function that it associates with every given set of feedback control policies. It will be shown that the online algorithm is mathematically equivalent to an offline iterative method, previously introduced in the literature, that solves the set of coupled algebraic Riccati equations (ARE) underlying the game problem using complete knowledge on the system dynamics. Here we show how the ADP techniques will enhance the capabilities of the offline method allowing an online solution without the requirement of complete knowledge of the system dynamics. The two participants in the continuous-time differential game are competing in real-time and the feedback Nash control strategies will be determined based on online measured data from the system. The algorithm is built on interplay between a learning phase, where each of the players is learning online the value that they associate with a given set of play policies, and a policy update step, performed by each of the payers towards decreasing the value of their cost. The players are learning concurrently. The feasibility of the ADP scheme is demonstrated in simulation.