XCS with Adaptive Action Mapping

XCS with Adaptive Action Mapping
复制标题

具有自适应动作映射的 XCS

DOI:
10.1007/978-3-642-34859-4_14
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
K. Takadama
K. Takadama
中科院分区:
--
文献类型:
--
作者:
Masaya Nakata;P. Lanzi;K. Takadama

文献摘要

被引文献

相似文献

XCS分类器系统进化出表示从状态-动作对到预期回报的完整映射的解决方案,因此,在任何可能的情况下,XCS都可以预测所有可用动作的值。这种完整的映射有时被认为是多余的,因为大多数应用程序(例如,分类)通常只关注最佳操作。在这篇文章中,我们介绍了XCS的一个扩展,它具有自适应(状态-动作)映射机制(或XCSAM),它进化出具有最大回报的以解决方案为中心的动作。虽然UCS发展的解决方案侧重于最佳可用操作,但只能解决监督分类问题,而我们的系统既可以解决监督问题,也可以解决多步骤问题,此外,它还可以根据问题调整映射的大小:最初,XCSAM开始构建完整的映射,然后慢慢尝试专注于可用的最佳操作。如果问题在每个小生境中只允许一个最优动作,那么随着进化的进行,XCSAM往往会专注于这样一个动作。如果有更多具有相同返回的操作可用,则XCSAM倾向于发展一个包含所有这些操作的映射。我们将XCSAM应用于有监督问题(布尔多路复用器)和多步迷宫类问题。我们的实验结果表明,XCSAM可以达到最优性能,但它需要的群体比XCS小,因为它进化的解决方案专注于每个子问题的最佳操作。
The XCS classifier system evolves solutions that represent complete mappings from state-action pairs to expected returns therefore, in every possible situation, XCS can predict the value of all the available actions. Such complete mapping is sometimes considered redundant as most of the applications (like for instance, classification), usually focus only on the best action. In this paper, we introduce an extension of XCS with an adaptive (state-action) mapping mechanism (or XCSAM) that evolves solutions focused actions with the largest returns. While UCS evolves solutions focused on the best available action but can only solve supervised classification problems, our system can solve both supervised and multi-step problems and, in addition, it can adapt the size of the mapping to the problems: Initially, XCSAM starts building a complete mapping and then it slowly tries to focus on the best actions available. If the problem admits only one optimal action in each niche, XCSAM tends to focus on such an action as the evolution proceeds. If more actions with the same return are available, XCSAM tends to evolve a mapping that includes all of them. We applied XCSAM both to supervised problems (the Boolean multiplexer) and to multi-step maze-like problems. Our experimental results show that XCSAM can reach optimal performance but requires smaller populations than XCS as it evolves solutions focused on the best actions available for each subproblem.