Position Control and Production of Various Strategies for Deep Learning Go Programs

Position Control and Production of Various Strategies for Deep Learning Go Programs
复制标题

DOI:
10.1109/taai48200.2019.8959895
复制
发表时间:
2019-11
期刊:
2019 International Conference on Technologies and Applications of Artificial Intelligence (TAAI)
影响因子:
--
通讯作者:
Tianwen Fan;Yuan Shi;Wanxiang Li;Kokolo Ikeda
Tianwen Fan;Yuan Shi;Wanxiang Li;Kokolo Ikeda
中科院分区:
其他
文献类型:
--
作者:
Tianwen Fan;Yuan Shi;Wanxiang Li;Kokolo Ikeda

文献摘要

相似文献

通过使用深度学习和强化学习技术,计算机围棋程序已经超过了顶级人类棋手。另一方面,“娱乐围棋AI”或“教练围棋AI”也是有趣的方向,没有得到很好的调查。为了娱乐初学者或中级玩家,已经进行了几项研究。位置控制或产生各种策略是重要的任务,已经提出了一些方法,并使用传统的蒙特卡罗树搜索程序进行了评估。在本文中,我们尝试将该方法应用于基于AlphaGo Zero的程序LeelaZero。以前的计划和新的计划之间有一些关键的区别。例如,新程序不使用随机模拟来结束游戏,那么以前产生各种策略的方法就不能使用了。在本文中,我们总结了差异和一些预期的问题,并提出了几种解决问题的方法。结果表明,改装后的LeelaZero可以温和地对抗实力较弱的球员(48%的人战胜了程序Ray)。通过对受试者的实验表明,平均每场比赛的非自然动作次数为1.22次,而不考虑自然度的简单方法平均为2.29次。我们还对所提出的训练“中锋”和“边缘/角球”球员的方法进行了评估,证实人类球员能够识别产生的策略(中锋或边缘/角球),概率为71.88%。
Computer Go programs have exceeded top-level human players by using deep learning and reinforcement learning techniques. On the other hand, “Entertainment Go AI” or “Coaching Go AI” are also interesting directions which have not been well investigated. Several researches have been done for entertaining beginners or intermediate players. Position control or producing various strategies are important tasks, and some methods have been proposed and evaluated using a traditional Monte-Carlo tree search program. In this paper, we try to adapt the method to LeelaZero, a program based on AlphaGo Zero. There are some critical differences between the previous program and the new program. For example the new program does not use random simulations to the ends of games, then the previous method for producing various strategies cannot be used. In this paper we summarized the differences and some expected problems, and proposed several approaches to solve the problems. It was shown that the modified LeelaZero could play gently against weaker players (48% won against a program Ray). Through experiments using human subjects, it was shown that the average number of unnatural moves per game was 1.22, where that by a simple method without considering naturalness was 2.29. Also we evaluated the proposed method for training “center-oriented” and “edge/corner-oriented” players, and it was confirmed that human players could identify the produced strategy (center or edge/corner) with a probability of 71.88%.