An Evolutionary Dynamical Analysis of Multi-Agent Learning in Iterated Games

An Evolutionary Dynamical Analysis of Multi-Agent Learning in Iterated Games
复制标题

DOI:
10.1007/s10458-005-3783-9
复制
发表时间:
2005
影响因子:
1.9
通讯作者:
K. Tuyls;P. Hoen;B. Vanschoenwinkel
K. Tuyls;P. Hoen;B. Vanschoenwinkel
中科院分区:
计算机科学4区
文献类型:
--
作者:
K. Tuyls;P. Hoen;B. Vanschoenwinkel

文献摘要

被引文献

相似文献

本文从进化动力学的角度研究了多智能体系统中的强化学习。典型的MAS是环境不是平稳的,马尔可夫性质是无效的。这就要求代理人具有适应性。RL是一种自然的方法来模拟单个代理的学习。然而,已知这些学习算法对单代理系统的参数设置的正确选择敏感。这个问题在MAS案例中更为普遍,因为代理之间的交互不断变化。这在很大程度上是一个开放的问题,MAS的开发人员如何设计的个人代理,通过学习,代理作为一个集体达到良好的解决方案。我们将表明,建模RL在MAS中,通过采取进化博弈论的观点,是一种新的和潜在的成功的方式来引导学习代理最适合的解决方案,他们手头的任务。我们展示了如何从进化博弈论的进化动力学(艾德)可以帮助开发人员的MAS在良好的选择所使用的RL算法的参数设置。艾德基本上预测MAS的均衡结果,其中代理使用单独的RL算法。更具体地说,我们展示了艾德如何预测迭代游戏的Q学习者的学习轨迹。此外,我们将我们的结果(的扩展)的集体智能框架(COIN)。COIN是MAS中协作任务学习的一种成熟的工程方法。代理人的效用被重新设计,以促进全球效用。我们展示了如何改进的结果MAS RL在COIN,和一个发达的扩展,预测的ED。
In this paper, we investigate Reinforcement learning (RL) in multi-agent systems (MAS) from an evolutionary dynamical perspective. Typical for a MAS is that the environment is not stationary and the Markov property is not valid. This requires agents to be adaptive. RL is a natural approach to model the learning of individual agents. These Learning algorithms are however known to be sensitive to the correct choice of parameter settings for single agent systems. This issue is more prevalent in the MAS case due to the changing interactions amongst the agents. It is largely an open question for a developer of MAS of how to design the individual agents such that, through learning, the agents as a collective arrive at good solutions. We will show that modeling RL in MAS, by taking an evolutionary game theoretic point of view, is a new and potentially successful way to guide learning agents to the most suitable solution for their task at hand. We show how evolutionary dynamics (ED) from Evolutionary Game Theory can help the developer of a MAS in good choices of parameter settings of the used RL algorithms. The ED essentially predict the equilibriums outcomes of the MAS where the agents use individual RL algorithms. More specifically, we show how the ED predict the learning trajectories of Q-Learners for iterated games. Moreover, we apply our results to (an extension of) the COllective INtelligence framework (COIN). COIN is a proved engineering approach for learning of cooperative tasks in MASs. The utilities of the agents are re-engineered to contribute to the global utility. We show how the improved results for MAS RL in COIN, and a developed extension, are predicted by the ED.