Comparison of rapid action value estimation variants for general game playing

Comparison of rapid action value estimation variants for general game playing
复制标题

一般游戏的快速动作值估计变体的比较

DOI:
--
复制
发表时间:
2016
期刊:
IEEE Conference on Computational Intelligence and Games
影响因子:
--
通讯作者:
M. Winands
M. Winands
中科院分区:
--
文献类型:
--
作者:
C. F. Sironi;M. Winands

文献摘要

参考文献

被引文献

相似文献

通用游戏(GGP)旨在创建能够在仅给定规则的情况下以专家级别玩任意游戏的计算机程序。由于缺乏特定于游戏的知识以及在线学习策略的必要性,使得蒙特卡洛树搜索(MCTS)成为应对 GGP 挑战的合适方法。有效的搜索控制机制可以显着提高 MCTS 的性能。 RAVE 策略及其更新的变体 GRAVE 就是出于这个原因而被提出的。在本文中,我们进一步研究了 GRAVE 在 GGP 中的使用,并将其性能与更成熟的 RAVE 策略以及使用更多全局信息的新变体(称为 HRAVE)进行比较。实验表明,对于某些游戏,GRAVE 和 HRAVE 的表现优于 RAVE,其中 GRAVE 是总体上最有前途的一种。
General Game Playing (GGP) aims at creating computer programs able to play any arbitrary game at an expert level given only its rules. The lack of game-specific knowledge and the necessity of learning a strategy online have made Monte-Carlo Tree Search (MCTS) a suitable method to tackle the challenges of GGP. An efficient search-control mechanism can substantially increase the performance of MCTS. The RAVE strategy and its more recent variant, GRAVE, have been proposed for this reason. In this paper we further investigate the use of GRAVE for GGP and compare its performance with the more established RAVE strategy and with a new variant, called HRAVE, that uses more global information. Experiments show that for some games GRAVE and HRAVE perform better than RAVE, with GRAVE being the most promising one overall.
一路强盗:UCB1 作为蒙特卡罗树搜索中的模拟策略
DOI: 10.1109/cig.2013.6633613
发表时间: 2013
期刊: --
影响因子: --
作者:
Powley E
通讯作者: Powley E