Mastering the game of Go without human knowledge

Mastering the game of Go without human knowledge
复制标题

DOI:
10.1038/nature24270
复制
发表时间:
2017-10-19
期刊:
影响因子:
64.8
通讯作者:
Hassabis, Demis
Hassabis, Demis
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Silver, David;Schrittwieser, Julian;Hassabis, Demis

文献摘要

被引文献

相似文献

人工智能的一个长期目标是一种算法,它可以学习,白板,在挑战性领域的超人能力。最近,AlphaGo成为第一个在围棋比赛中击败世界冠军的程序。AlphaGo中的树搜索使用深度神经网络评估位置和选择的移动。这些神经网络通过人类专家动作的监督学习和自我游戏的强化学习进行训练。在这里,我们介绍了一种完全基于强化学习的算法,没有人类数据,指导或游戏规则之外的领域知识。AlphaGo成为自己的老师:一个神经网络被训练来预测AlphaGo自己的走法选择,以及AlphaGo游戏的赢家。该神经网络提高了树搜索的强度,从而在下一次迭代中产生更高质量的移动选择和更强的自我发挥。从白板开始,我们的新程序AlphaGo Zero实现了超人的表现,以100比0战胜了之前公布的冠军AlphaGo。
A long-standing goal of artificial intelligence is an algorithm that learns, tabula rasa, superhuman proficiency in challenging domains. Recently, AlphaGo became the first program to defeat a world champion in the game of Go. The tree search in AlphaGo evaluated positions and selected moves using deep neural networks. These neural networks were trained by supervised learning from human expert moves, and by reinforcement learning from self-play. Here we introduce an algorithm based solely on reinforcement learning, without human data, guidance or domain knowledge beyond game rules. AlphaGo becomes its own teacher: a neural network is trained to predict AlphaGo's own move selections and also the winner of AlphaGo's games. This neural network improves the strength of the tree search, resulting in higher quality move selection and stronger self-play in the next iteration. Starting tabula rasa, our new program AlphaGo Zero achieved superhuman performance, winning 100-0 against the previously published, champion-defeating AlphaGo.