Nash Q-learning for general-sum stochastic games

Nash Q-learning for general-sum stochastic games
复制标题

DOI:
10.1162/1532443041827880
复制
发表时间:
2004-08-15
影响因子:
6
通讯作者:
Wellman, MP
Wellman, MP
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hu, JL;Wellman, MP

文献摘要

被引文献

相似文献

我们扩展Q-学习的非合作多智能体的情况下,使用的框架一般和随机游戏。学习代理在联合动作上维护Q函数,并基于假设在当前Q值上的纳什均衡行为来执行更新。这个学习协议可证明收敛给定的阶段游戏(由Q值定义),在学习过程中出现的某些限制。一对两人网格游戏的实验表明,这种对游戏结构的限制并不一定是必需的。在两种网格环境中学习时遇到的阶段游戏都违反了条件。然而,学习在第一个网格游戏中始终收敛,它具有唯一的均衡Q函数,但有时在第二个网格游戏中无法收敛,它具有三个不同的均衡Q函数。在两个游戏的离线学习性能的比较中,我们发现代理更有可能达到联合最优路径与纳什Q-学习比单代理Q-学习方法。当至少一个智能体采用纳什Q学习时,两个智能体的性能优于使用单智能体Q学习。我们还实现了一个在线版本的纳什Q学习,平衡了探索与利用,提高了性能。
We extend Q-learning to a noncooperative multiagent context, using the framework of general-sum stochastic games. A learning agent maintains Q-functions over joint actions, and performs updates based on assuming Nash equilibrium behavior over the current Q-values. This learning protocol provably converges given certain restrictions on the stage games (defined by Q-values) that arise during learning. Experiments with a pair of two-player grid games suggest that such restrictions on the game structure are not necessarily required. Stage games encountered during learning in both grid environments violate the conditions. However, learning consistently converges in the first grid game, which has a unique equilibrium Q-function, but sometimes fails to converge in the second, which has three different equilibrium Q-functions. In a comparison of offline learning performance in both games, we find agents are more likely to reach a joint optimal path with Nash Q-learning than with a single-agent Q-learning method. When at least one agent adopts Nash Q-learning, the performance of both agents is better than using single-agent Q-learning. We have also implemented an online version of Nash Q-learning that balances exploration with exploitation, yielding improved performance.