Monte Carlo Tree Search with Robust Exploration
Monte Carlo Tree Search with Robust Exploration
复制标题
具有稳健探索的蒙特卡罗树搜索
DOI:
10.1007/978-3-319-50935-8_4
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
T. Imagawa and T. Kaneko
中科院分区:
文献类型:
--
作者:
山田亮太;松谷悠佑;木村尭朗;伊達広行;Takahisa Imagawa and Tomoyuki Kaneko;今川 孝久,金子知適;T. Imagawa and T. Kaneko
This paper presents a new Monte-Carlo tree search method that focuses on identifying the best move. UCT which minimizes the cumulative regret, has achieved remarkable success in Go and other games. However, recent studies on simple regret reveal that there are better exploration strategies. To further improve the performance, a leaf to be explored is determined not only by the mean but also by the whole reward distribution. We adopted a hybrid approach to obtain reliable distributions. A negamax-style backup of reward distributions is used in the shallower half of a search tree, and UCT is adopted in the rest of the tree. Experiments on synthetic trees show that this presented method outperformed UCT and similar methods, except for trees having uniform width and depth.