The Resilience of Cooperation in a Dilemma Game Played by Reinforcement Learning Agents

The Resilience of Cooperation in a Dilemma Game Played by Reinforcement Learning Agents
复制标题

强化学习代理在困境博弈中的合作弹性

DOI:
10.1109/agents.2017.8015297
复制
发表时间:
2017
期刊:
Proceedings of the 2nd IEEE International Conference on Agents
影响因子:
--
通讯作者:
and Nobuhiro Inuzuka
and Nobuhiro Inuzuka
中科院分区:
--
文献类型:
--
作者:
Koichi Moriyama;Kaori Nakase;Atsuko Mutoh;and Nobuhiro Inuzuka

文献摘要

相似文献

这项工作讨论了(独立的)强化学习代理可以在多代理环境中做什么。特别是,我们考虑一个无状态的Q学习代理的囚徒困境(PD)的游戏。虽然在文献中已经表明,无状态的,独立的Q-学习代理已经很难相互合作的迭代PD(IPD)游戏,我们给了PD支付和Q-学习参数,帮助代理相互合作的条件。在此基础上,我们还讨论了IPD博弈中双方合作的概率。它认为相互合作是脆弱的,即,一次不幸的背叛会让特工们滑向相互背叛的螺旋。然而,它并不总是正确的。相互合作将加强自身,因此将是强有力的和有弹性的。因此,这项工作分析得出多久一系列的相互合作,一旦它发生,同时考虑到弹性。它使我们进一步理解IPD游戏中的强化学习过程。
This work discusses what an (independent) reinforcement learning agent can do in a multiagent environment. In particular, we consider a stateless Q-learning agent in a Prisoner's Dilemma (PD) game. Although it had been shown in the literature that stateless, independent Q-learning agents had been difficult to cooperate with each other in an iterated PD (IPD) game, we gave a condition of PD payoffs and Q-learning parameters that helps the agents cooperate with each other. Based on the condition, we also discussed the ratio of mutual cooperation happening in IPD games. It supposed that mutual cooperation was fragile, i.e., one misfortune defection would have the agents slide down the spiral of mutual defection. However, it is not always correct. Mutual cooperation will reinforce itself and thus it will be robust and resilient. Hence, this work analytically derives how long a series of mutual cooperation continues once it happened while considering the resilience. It gives us further comprehension of the process of reinforcement learning in IPD games.