An Alternative Multitask Training for Evaluation Functions in the Game of Go

An Alternative Multitask Training for Evaluation Functions in the Game of Go
复制标题

围棋评估函数的另一种多任务训练

DOI:
10.1109/taai.2018.00037
复制
发表时间:
2018
期刊:
IEEE Technologies and Applications of Artificial Intelligence
影响因子:
--
通讯作者:
Yusaku Mandai and Tomoyuki Kaneko
Yusaku Mandai and Tomoyuki Kaneko
中科院分区:
--
文献类型:
--
作者:
Hiroki Tamari;Shohei Nakamura;Shigeru Takano;Yoshihiro Okada;田中匠,武田直人,関洋平;Yusaku Mandai and Tomoyuki Kaneko

文献摘要

相似文献

对于围棋、国际象棋和将棋(日本象棋)游戏,深度神经网络 (DNN) 有助于构建准确的评估函数,并且许多研究尝试创建所谓的价值网络,用于预测给定状态的奖励。最近对围棋游戏价值网络的研究表明,具有两个不同目标的双头神经网络可以有效地进行训练,并且比单头网络表现更好。两个头之一称为价值头,另一个称为政策头,预测给定状态的下一步行动。这种多任务训练使网络更加鲁棒并提高泛化性能。在本文中,我们证明简单的判别器网络是多任务学习的替代目标。与现有的深度神经网络相比,我们提出的网络由于其简单的输出而可以更容易地设计。我们的实验结果表明,我们的判别目标也使学习变得稳定,并且我们的方法训练的评估函数在预测下一步行动和比赛强度方面与现有研究的训练相当。
For the game of Go, Chess, and Shogi (Japanese Chess), deep neural networks (DNNs) have contributed to building accurate evaluation functions, and many studies have attempted to create the so-called value network, which predicts the reward of a given state. A recent study of the value network for the game of Go has shown that a two-headed neural network with two different objectives can be trained effectively and performs better than a single-headed network. One of the two heads is called a value head and the other head, the policy head, predicts the next move at a given state. This multitask training makes the network more robust and improves the generalization performance. In this paper, we show that a simple discriminator network is an alternative target of multitask learning. Compared to the existing deep neural network, our proposed network can be designed more easily because of its simple output. Our experimental results showed that our discriminative target also makes the learning stable and the evaluation function trained by our method is comparable to the training of existing studies in terms of predicting the next move and playing strength.