Recurrent Experience Replay in Distributed Reinforcement Learning

Recurrent Experience Replay in Distributed Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Steven Kapturowski;Georg Ostrovski;John Quan;R. Munos;Will Dabney
Steven Kapturowski;Georg Ostrovski;John Quan;R. Munos;Will Dabney
中科院分区:
其他
文献类型:
--
作者:
Steven Kapturowski;Georg Ostrovski;John Quan;R. Munos;Will Dabney

文献摘要

被引文献

相似文献

在最近成功的分布式训练RL智能体的基础上,本文研究了基于rnn的分布式优先体验重放训练RL智能体。我们研究了参数滞后对表征漂移和循环状态失效的影响,并通过经验推导出一种改进的训练策略。使用单一网络架构和固定的超参数集,最终生成的代理,Recurrent Replay Distributed DQN,是Atari-57上先前技术水平的四倍,并与DMLab-30上的技术水平相匹配。在57款雅达利游戏中,它是第一个在52款游戏中超越人类水平的智能体。
Building on the recent successes of distributed training of RL agents, in this paper we investigate the training of RNN-based RL agents from distributed prioritized experience replay. We study the effects of parameter lag resulting in representational drift and recurrent state staleness and empirically derive an improved training strategy. Using a single network architecture and fixed set of hyper-parameters, the resulting agent, Recurrent Replay Distributed DQN, quadruples the previous state of the art on Atari-57, and matches the state of the art on DMLab-30. It is the first agent to exceed human-level performance in 52 of the 57 Atari games.