SAI: a Sensible Artificial Intelligence that plays with handicap and targets high scores in 9x9 Go (extended version)

SAI: a Sensible Artificial Intelligence that plays with handicap and targets high scores in 9x9 Go (extended version)
复制标题

SAI:一种明智的人工智能,可以在 9x9 围棋中克服让分并以高分为目标(扩展版)

DOI:
--
复制
发表时间:
2019
期刊:
European Conference on Artificial Intelligence
影响因子:
--
通讯作者:
M. Parton
M. Parton
中科院分区:
--
文献类型:
--
作者:
F. Morandin;G. Amato;M. Fantozzi;R. Gini;C. Metta;M. Parton

文献摘要

被引文献

相似文献

我们开发了一种新模型,可以应用于任何完美信息两人零和博弈,以取得高分,从而获得完美的比赛。我们将此模型集成到 Google DeepMind 与 AlphaGo 引入的蒙特卡罗树搜索策略迭代学习管道中。在 9x9 Go 上训练这个模型会产生一个超人的围棋棋手,从而证明它是稳定和鲁棒的。我们表明,该模型可用于有效地应对位置和得分障碍,并最大限度地减少次优动作。我们开发了一系列代理,可以针对任何对手取得高分,并在对抗弱对手时从非常严重的劣势中恢复过来。据我们所知,这些是该方向的第一个有效成果。
We develop a new model that can be applied to any perfect information two-player zero-sum game to target a high score, and thus a perfect play. We integrate this model into the Monte Carlo tree search-policy iteration learning pipeline introduced by Google DeepMind with AlphaGo. Training this model on 9x9 Go produces a superhuman Go player, thus proving that it is stable and robust. We show that this model can be used to effectively play with both positional and score handicap, and to minimize suboptimal moves. We develop a family of agents that can target high scores against any opponent, and recover from very severe disadvantage against weak opponents. To the best of our knowledge, these are the first effective achievements in this direction.