Ieee Transactions on Computational Intelligence and Ai in Games 1 N-grams and the Last-good-reply Policy Applied in General Game Playing

Ieee Transactions on Computational Intelligence and Ai in Games 1 N-grams and the Last-good-reply Policy Applied in General Game Playing
复制标题

Ieee Transactions on Computational Intelligence and Ai in Games 1 N-grams 和一般游戏中应用的最后良好回复策略

DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
Y. Björnsson
Y. Björnsson
中科院分区:
--
文献类型:
--
作者:
Mandy J. W. Tak;M. Winands;Y. Björnsson

文献摘要

被引文献

相似文献

通用游戏(GGP)的目标是创建能够在专家级别玩各种不同游戏的程序,只考虑游戏规则。最成功的GGP程序目前采用基于模拟的蒙特-卡罗树搜索(MCTS)。MCTS的性能在很大程度上取决于所使用的仿真策略。在本文中,我们介绍了改进的GGP模拟策略,我们实现和测试的GGP代理CADIAPLAYER,赢得了国际GGP比赛在2007年和2008年。有两个方面的改进:首先,我们表明,一个简单的贪婪的探索策略更好地工作在模拟播放比softmax为基础的吉布斯测量目前使用的CADIAPLAYER,其次,我们介绍了一个通用的框架,基于N-Gram学习有前途的移动序列。总的来说,这些增强功能大大提高了CADIAPLAYER的性能。例如,在我们的测试套件中,包括五种不同的双人回合制游戏,它们的平均胜率约为70%。这些增强功能在多人游戏和错误移动游戏中也很有效。我们还进行了实验与最后一个好的答复政策。最后一个好的答复政策结合N-Gram也进行了测试。最后一次好答复策略已经在Go程序中被证明是成功的,我们证明它在GGP中也有希望。
—The aim of General Game Playing (GGP) is to create programs capable of playing a wide range of different games at an expert level, given only the rules of the game. The most successful GGP programs currently employ simulation-based Monte-Carlo Tree Search (MCTS). The performance of MCTS depends heavily on the simulation strategy used. In this paper we introduce improved simulation strategies for GGP that we implement and test in the GGP agent CADIAPLAYER, which won the International GGP competition in both 2007 and 2008. There are two aspects to the improvements: first, we show that a simple ϵ-greedy exploration strategy works better in the simulation play-outs than the softmax-based Gibbs measure currently used in CADIAPLAYER and, secondly, we introduce a general framework based on N-Grams for learning promising move sequences. Collectively, these enhancements result in a much improved performance of CADIAPLAYER. For example, in our test-suite consisting of five different two-player turn-based games, they led to an impressive average win rate of approximately 70%. The enhancements are also shown to be effective in multi-player and simultaneous-move games. We additionally perform experiments with the Last-Good-Reply Policy. The Last-Good-Reply Policy combined with N-Grams is also tested. The Last-Good-Reply Policy has already been shown to be successful in Go programs and we demonstrate that it also has promise in GGP.