The Power of Forgetting: Improving the Last-Good-Reply Policy in Monte Carlo Go

The Power of Forgetting: Improving the Last-Good-Reply Policy in Monte Carlo Go
复制标题

遗忘的力量:改进 Monte Carlo Go 中的最后良好回复策略

DOI:
--
复制
发表时间:
2010
影响因子:
--
通讯作者:
P. Drake
P. Drake
中科院分区:
工程技术4区
文献类型:
--
作者:
Hendrik Baier;P. Drake

文献摘要

被引文献

相似文献

演奏GO游戏的程序的主要范例是Monte Carlo Tree搜索。该算法通过玩许多模拟游戏(播放)来构建搜索树。每个竞赛都由树内的一系列动作组成,然后是树超越树的许多动作。超越树的移动是由有偏见的随机抽样策略生成的。最近发表的最后一项善良的政策采取了行动,在先前的竞争中,它已经成功地回复了直接前进的举动。本文提出了对这项政策的修改,不仅纪念最近成功的动作,而且立即忘记了最近失败的举动。这种修改为比赛实力提供了很大的改善。我们还表明,对前两个动作的响应优于回应前一个动作。令人惊讶的是,记住每个答复的胜利率比仅仅记得最后一个好的答复(实际上比不存储好的答复更糟糕)要差得多。
The dominant paradigm for programs playing the game of Go is Monte Carlo tree search. This algorithm builds a search tree by playing many simulated games (playouts). Each playout consists of a sequence of moves within the tree followed by many moves beyond the tree. Moves beyond the tree are generated by a biased random sampling policy. The recently published last-good-reply policy makes moves that, in previous playouts, have been successful replies to immediately preceding moves. This paper presents a modification of this policy that not only remembers moves that recently succeeded but also immediately forgets moves that recently failed. This modification provides a large improvement in playing strength. We also show that responding to the previous two moves is superior to responding to the previous one move. Surprisingly, remembering the win rate of every reply performs much worse than simply remembering the last good reply (and indeed worse than not storing good replies at all).