Evaluation of Loss Function for Stable Policy Learning in Dobutsu Shogi
Evaluation of Loss Function for Stable Policy Learning in Dobutsu Shogi
复制标题
DOI:
10.1109/taai51410.2020.00044
复制
发表时间:
2020-12
期刊:
影响因子:
--
通讯作者:
T. Nakayashiki;Tomoyuki Kaneko
中科院分区:
文献类型:
--
作者:
T. Nakayashiki;Tomoyuki Kaneko
AlphaZero achieved superhuman performance in Go, chess, and shogi, supported by massive computational re-sources. This paper investigates policy learning in AlphaZero, aiming at improving its learning efficiency. AlphaZero trains policy and value networks that are similar to those in a standard reinforcement learning setting, but in a different way closely coupled with game tree search in self-play. Inspired by recent studies that investigate policy optimization in tree search, we apply PPO, which is a state-of-the method for learning policy in reinforcement learning, into AlphaZero’s style learning. Then, we show its effectiveness in a small but interesting game where the database made by retrograde analysis is available. The results support that our method improves the sample efficiency in the learning.