Analyses of Tabular AlphaZero on Strongly-Solved Stochastic Games

Analyses of Tabular AlphaZero on Strongly-Solved Stochastic Games
复制标题

DOI:
10.1109/access.2023.3246638
复制
发表时间:
2023
期刊:
影响因子:
3.9
通讯作者:
Chu-Hsuan Hsueh;Kokolo Ikeda;I-Chen Wu;Jr-Chang Chen;T. Hsu
Chu-Hsuan Hsueh;Kokolo Ikeda;I-Chen Wu;Jr-Chang Chen;T. Hsu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chu-Hsuan Hsueh;Kokolo Ikeda;I-Chen Wu;Jr-Chang Chen;T. Hsu

文献摘要

相似文献

AlphaZero算法在国际象棋、手势和围棋中通过学习达到超人水平,除了游戏规则外,不需要特定领域的知识。本文以随机博弈为研究对象,研究AlphaZero能否学习理论值和最优对策。由于随机博弈的理论值是预期胜率,而不是简单的赢、输或平局,因此值得研究AlphaZero近似头寸的预期胜率的能力。本文还深入研究了超参数对AlphaZero的影响以及一些实现细节。分析主要基于AlphaZero学习和查找表。深度神经网络(DNN)与原始AlphaZero中的网络也进行了实验和比较。测试的随机游戏包括中国黑暗象棋和爱因斯坦乌菲特!的简化和强解变体。实验表明,AlphaZero可以学习对最佳玩家几乎是最优的策略,并可以准确地学习价值。更详细地说,这样好的结果是通过在较大范围内设置不同的超参数来实现的,尽管观察到规模较大的游戏往往具有略窄的适当超参数范围。此外,使用DNN学习的结果与查找表相似。
The AlphaZero algorithm achieved superhuman levels of play in chess, shogi, and Go by learning without domain-specific knowledge except for game rules. This paper targets stochastic games and investigates whether AlphaZero can learn theoretical values and optimal play. Since the theoretical values of stochastic games are expected win rates, not a simple win, loss, or draw, it is worth investigating the ability of AlphaZero to approximate expected win rates of positions. This paper also thoroughly studies how AlphaZero is influenced by hyper-parameters and some implementation details. The analyses are mainly based on AlphaZero learning with lookup tables. Deep neural networks (DNNs) like the ones in the original AlphaZero are also experimented and compared. The tested stochastic games include reduced and strongly-solved variants of Chinese dark chess and EinStein würfelt nicht!. The experiments showed that AlphaZero could learn policies that play almost optimally against the optimal player and could learn values accurately. In more detail, such good results were achieved by different hyper-parameter settings in a wide range, though it was observed that games on larger scales tended to have a little narrower range of proper hyper-parameters. In addition, the results of learning with DNNs were similar to lookup tables.