Numerical analysis of a reinforcement learning model with the dynamic aspiration level in the iterated Prisoner's dilemma

Numerical analysis of a reinforcement learning model with the dynamic aspiration level in the iterated Prisoner's dilemma
复制标题

DOI:
10.1016/j.jtbi.2011.03.005
复制
发表时间:
2011-06-07
影响因子:
2
通讯作者:
Nakamura, Mitsuhiro
Nakamura, Mitsuhiro
中科院分区:
生物学4区
文献类型:
--
作者:
Masuda, Naoki;Nakamura, Mitsuhiro

文献摘要

被引文献

相似文献

人类和其他动物可以根据环境线索(包括通过经验获得的反馈)调整其社会行为。然而,基于经验的学习的球员在进化和社会困境游戏中的合作维持的效果仍然相对不清楚。一些先前的文献表明,学习的球员的相互合作是困难的,或者需要一个复杂的学习模型。在迭代囚徒困境的背景下,我们数值研究的强化学习模型的性能。我们的模型修改了Karandikar et al.(1998),Posch et al.(1999),Macy and Flache(2002)的模型,其中如果获得的回报大于动态阈值,则参与者满意。我们发现,如果阈值的动态不是太快,强化信号和下一轮动作之间的关联足够强,遵守修改后的学习的球员相互合作的概率很高。学习型玩家也能有效地对抗反应型策略。在进化动力学中,它们可以入侵采用简单但有竞争力的策略的玩家群体。我们版本的强化学习模型不会使以前的模型复杂化,并且足够简单但灵活。它可能有助于探索学习和进化之间的关系,在社会困境的情况下。(C)2011爱思唯尔有限公司保留所有权利。
Humans and other animals can adapt their social behavior in response to environmental cues including the feedback obtained through experience. Nevertheless, the effects of the experience-based learning of players in evolution and maintenance of cooperation in social dilemma games remain relatively unclear. Some previous literature showed that mutual cooperation of learning players is difficult or requires a sophisticated learning model. In the context of the iterated Prisoner's dilemma, we numerically examine the performance of a reinforcement learning model. Our model modifies those of Karandikar et al. (1998), Posch et al. (1999), and Macy and Flache (2002) in which players satisfice if the obtained payoff is larger than a dynamic threshold. We show that players obeying the modified learning mutually cooperate with high probability if the dynamics of threshold is not too fast and the association between the reinforcement signal and the action in the next round is sufficiently strong. The learning players also perform efficiently against the reactive strategy. In evolutionary dynamics, they can invade a population of players adopting simpler but competitive strategies. Our version of the reinforcement learning model does not complicate the previous model and is sufficiently simple yet flexible. It may serve to explore the relationships between learning and evolution in social dilemma situations. (C) 2011 Elsevier Ltd. All rights reserved.