An examination of evolved behavior in two reinforcement learning systems

An examination of evolved behavior in two reinforcement learning systems
复制标题

DOI:
10.1016/j.dss.2013.01.019
复制
发表时间:
2013-04
期刊:
Decis. Support Syst.
影响因子:
--
通讯作者:
D. A. Gaines;Ramakrishnan Pakath
D. A. Gaines;Ramakrishnan Pakath
中科院分区:
其他
文献类型:
--
作者:
D. A. Gaines;Ramakrishnan Pakath

文献摘要

被引文献

相似文献

利用基于智能体的仿真实验,我们评估了两种强化学习系统(RLS)范式——经典的学习分类器系统(LCS)和一种增强的扩展分类器系统(XCS)——在迭代囚徒困境(IPD)博弈中的相对性能。在先前的研究中,XCS在解决动物与迷宫和布尔多路复用测试问题上优于LCS。我们的工作与这些努力有重叠,并且是这些努力的延伸,因为它允许评估每个系统的能力(a)应对延迟的环境反馈,(b)将非理性选择进化为最佳行为,以及(c)应对来自环境的不可预测的输入。我们发现,虽然XCS在四个关键性能指标上明显优于LCS,但在与确定性的、反应性的游戏代理(针锋相对)进行IPD游戏时,LCS在与不可预测的对手(Rand)进行游戏时表现更好,尽管需要大量的进化努力。此外,在单独检查每个XCS增强后,我们看到配备单个XCS功能的特定LCS变体在针对两种类型对手的特定指标方面比传统LCS模型和/或XCS模型做得更好,但通常需要更大的进化努力。这表明,如果离线(而不是在线)关注性能和特定性能目标,那么可以构建相对简单的LCS变体,而不是成熟的XCS系统。使用配备XCS功能组合的LCS变体进行进一步评估,将有助于更好地理解这些功能对IPD性能的协同影响。
Using agent-based simulation experiments, we assess the relative performance of two Reinforcement Learning System (RLS) paradigms – the classical Learning Classifier System (LCS) and an enhancement, the Extended Classifier System (XCS) – in the context of playing the Iterated Prisoner's Dilemma (IPD) game. In prior research, the XCS outperforms the LCS in solving the Animats-and-Maze and Boolean Multiplexer test problems. Our work has overlaps with and is an extension of such efforts in that it allows assessment of each system's ability to (a) cope with delayed environmental feedback, (b) evolve irrational choice as the optimal behavior, and (c) cope with unpredictable input from the environment. We find that while the XCS is considerably superior to the LCS, in terms of four key performance metrics, in playing IPD games against a deterministic, reactive game-playing agent (Tit-for-Tat), the LCS does better against an unpredictable opponent (Rand) albeit with significant evolutionary effort. Further, upon examining each XCS enhancement in isolation, we see that specific LCS variants equipped with a single XCS feature, do better than the traditional LCS model and/or the XCS model in terms of particular metrics against both types of opponents but, again, usually with greater evolutionary effort. This suggests that if offline, rather than online, performance and specific performance goals are the focus, then one may construct relatively-simpler LCS variants rather than full-fledged XCS systems. Further assessments using LCS variants equipped with combinations of XCS features should help better comprehend the synergistic impacts of these features on performance in the IPD.