Multiple Policy Value Monte Carlo Tree Search

Multiple Policy Value Monte Carlo Tree Search
复制标题

DOI:
10.24963/ijcai.2019/653
复制
发表时间:
2019-05
期刊:
--
影响因子:
--
通讯作者:
Li-Cheng Lan;Wei Li;Ting-Han Wei;I-Chen Wu
Li-Cheng Lan;Wei Li;Ting-Han Wei;I-Chen Wu
中科院分区:
其他
文献类型:
--
作者:
Li-Cheng Lan;Wei Li;Ting-Han Wei;I-Chen Wu

文献摘要

被引文献

相似文献

许多最强的游戏程序使用蒙特卡罗树搜索(MCTS)和深度神经网络(DNN)的组合,其中DNN被用作策略或价值评估器。考虑到有限的预算,例如在线游戏或在AlphaZero (AZ)训练的自我游戏阶段,需要在准确的状态估计和更多的MCTS模拟之间达到平衡,这两者对于强大的游戏代理都是至关重要的。通常,较大的dnn在泛化和准确评估方面表现更好,而较小的dnn成本较低,因此可以在相同的预算下进行更多的MCTS模拟和更大的搜索树。本文介绍了一种新的方法,即多策略值神经网络(MPV-MCTS),它将不同规模的多个策略值神经网络(pv - nn)组合在一起,以保持每个网络的优势,其中使用两个pv - nn f_S和f_L。我们通过游戏NoGo的实验表明,f_S和f_L组合的MPV-MCTS优于具有策略值MCTS的单个PV-NN,称为PV-MCTS。此外,MPV-MCTS在AZ训练方面也优于PV-MCTS。
Many of the strongest game playing programs use a combination of Monte Carlo tree search (MCTS) and deep neural networks (DNN), where the DNNs are used as policy or value evaluators. Given a limited budget, such as online playing or during the self-play phase of AlphaZero (AZ) training, a balance needs to be reached between accurate state estimation and more MCTS simulations, both of which are critical for a strong game playing agent. Typically, larger DNNs are better at generalization and accurate evaluation, while smaller DNNs are less costly, and therefore can lead to more MCTS simulations and bigger search trees with the same budget. This paper introduces a new method called the multiple policy value MCTS (MPV-MCTS), which combines multiple policy value neural networks (PV-NNs) of various sizes to retain advantages of each network, where two PV-NNs f_S and f_L are used in this paper. We show through experiments on the game NoGo that a combined f_S and f_L MPV-MCTS outperforms single PV-NN with policy value MCTS, called PV-MCTS. Additionally, MPV-MCTS also outperforms PV-MCTS for AZ training.