Minigo: A Case Study in Reproducing Reinforcement Learning Research
Minigo: A Case Study in Reproducing Reinforcement Learning Research
复制标题
DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Brian Lee;Andrew Jackson;T. Madams;Seth Troisi;Derek Jones
中科院分区:
文献类型:
--
作者:
Brian Lee;Andrew Jackson;T. Madams;Seth Troisi;Derek Jones
The reproducibility of reinforcement-learning research has been highlighted as a key challenge area in the field. In this paper, we present a case study in reproducing the results of one groundbreaking algorithm, AlphaZero, a reinforcement learning system that learns how to play Go at a superhuman level given only the rules of the game. We describe Minigo, a reproduction of the AlphaZero system using publicly available Google Cloud Platform infrastructure and Google Cloud TPUs. The Minigo system includes both the central reinforcement learning loop as well as auxiliary monitoring and evaluation infrastructure. With ten days of training from scratch on 800 Cloud TPUs, Minigo can play evenly against LeelaZero and ELF OpenGo, two of the strongest publicly available Go AIs. We discuss the difficulties of scaling a reinforcement learning system and the monitoring systems required to understand the complex interplay of hyperparameter configurations.