Minigo: A Case Study in Reproducing Reinforcement Learning Research

Minigo: A Case Study in Reproducing Reinforcement Learning Research
复制标题

DOI:
--
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Brian Lee;Andrew Jackson;T. Madams;Seth Troisi;Derek Jones
Brian Lee;Andrew Jackson;T. Madams;Seth Troisi;Derek Jones
中科院分区:
其他
文献类型:
--
作者:
Brian Lee;Andrew Jackson;T. Madams;Seth Troisi;Derek Jones

文献摘要

被引文献

相似文献

强化学习研究的可重复性已成为该领域的一个关键挑战。在本文中,我们提出了一个案例研究,再现了一个突破性算法AlphaZero的结果,AlphaZero是一个强化学习系统,它只在给定游戏规则的情况下学习如何以超人的水平下围棋。我们描述了Minigo, AlphaZero系统的复制,使用公开可用的谷歌云平台基础设施和谷歌云tpu。Minigo系统既包括中央强化学习回路,也包括辅助监测和评估基础设施。在800个Cloud tpu上从零开始训练10天后,Minigo可以与LeelaZero和ELF OpenGo这两种最强大的公开围棋人工智能平起平坐。我们讨论了扩展强化学习系统和监控系统所需的困难,以理解超参数配置的复杂相互作用。
The reproducibility of reinforcement-learning research has been highlighted as a key challenge area in the field. In this paper, we present a case study in reproducing the results of one groundbreaking algorithm, AlphaZero, a reinforcement learning system that learns how to play Go at a superhuman level given only the rules of the game. We describe Minigo, a reproduction of the AlphaZero system using publicly available Google Cloud Platform infrastructure and Google Cloud TPUs. The Minigo system includes both the central reinforcement learning loop as well as auxiliary monitoring and evaluation infrastructure. With ten days of training from scratch on 800 Cloud TPUs, Minigo can play evenly against LeelaZero and ELF OpenGo, two of the strongest publicly available Go AIs. We discuss the difficulties of scaling a reinforcement learning system and the monitoring systems required to understand the complex interplay of hyperparameter configurations.