Evaluation of Loss Function for Stable Policy Learning in Dobutsu Shogi

Evaluation of Loss Function for Stable Policy Learning in Dobutsu Shogi
复制标题

DOI:
10.1109/taai51410.2020.00044
复制
发表时间:
2020-12
期刊:
2020 International Conference on Technologies and Applications of Artificial Intelligence (TAAI)
影响因子:
--
通讯作者:
T. Nakayashiki;Tomoyuki Kaneko
T. Nakayashiki;Tomoyuki Kaneko
中科院分区:
其他
文献类型:
--
作者:
T. Nakayashiki;Tomoyuki Kaneko

文献摘要

相似文献

AlphaZero在围棋、国际象棋和将棋方面取得了超人的表现,并得到了大量计算资源的支持。本文研究了AlphaZero的策略学习,旨在提高其学习效率。AlphaZero训练的策略和价值网络类似于标准强化学习设置中的策略和价值网络,但以不同的方式与自我游戏中的游戏树搜索紧密结合。受最近研究树搜索中策略优化的启发,我们将PPO(强化学习中学习策略的最新方法)应用到AlphaZero的风格学习中。然后,我们在一个小而有趣的游戏中展示了它的有效性,其中通过逆行分析得到的数据库可用。实验结果表明,该方法提高了学习过程中的样本效率。
AlphaZero achieved superhuman performance in Go, chess, and shogi, supported by massive computational re-sources. This paper investigates policy learning in AlphaZero, aiming at improving its learning efficiency. AlphaZero trains policy and value networks that are similar to those in a standard reinforcement learning setting, but in a different way closely coupled with game tree search in self-play. Inspired by recent studies that investigate policy optimization in tree search, we apply PPO, which is a state-of-the method for learning policy in reinforcement learning, into AlphaZero’s style learning. Then, we show its effectiveness in a small but interesting game where the database made by retrograde analysis is available. The results support that our method improves the sample efficiency in the learning.