The Dynamics of Reinforcement Learning in Cooperative Multiagent Systems

The Dynamics of Reinforcement Learning in Cooperative Multiagent Systems
复制标题

DOI:
--
复制
发表时间:
1998-07
期刊:
--
影响因子:
--
通讯作者:
C. Claus;Craig Boutilier
C. Claus;Craig Boutilier
中科院分区:
其他
文献类型:
--
作者:
C. Claus;Craig Boutilier

文献摘要

被引文献

相似文献

强化学习可以为多智能体系统中的智能体提供一种鲁棒而自然的方法来学习如何协调它们的动作选择。我们研究的一些因素,可以影响动态的学习过程中,这样的设置。我们首先区分不知道(或忽视)其他代理存在的强化学习者和那些明确试图学习联合行动的价值及其对手策略的强化学习者。我们研究(一个简单的形式)Q-学习在合作多智能体系统在这两个角度下,专注于游戏结构和探索策略的影响收敛到(最优和次优)纳什均衡。然后,我们提出了替代乐观的探索策略,增加收敛到最优均衡的可能性。
Reinforcement learning can provide a robust and natural means for agents to learn how to coordinate their action choices in multi agent systems. We examine some of the factors that can influence the dynamics of the learning process in such a setting. We first distinguish reinforcement learners that are unaware of (or ignore) the presence of other agents from those that explicitly attempt to learn the value of joint actions and the strategies of their counterparts. We study (a simple form of) Q-leaming in cooperative multi agent systems under these two perspectives, focusing on the influence of that game structure and exploration strategies on convergence to (optimal and suboptimal) Nash equilibria. We then propose alternative optimistic exploration strategies that increase the likelihood of convergence to an optimal equilibrium.