Multi-task Deep Reinforcement Learning with PopArt

Multi-task Deep Reinforcement Learning with PopArt
复制标题

DOI:
10.1609/aaai.v33i01.33013796
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Matteo Hessel;Hubert Soyer;L. Espeholt;Wojciech M. Czarnecki;Simon Schmitt;H. V. Hasselt
Matteo Hessel;Hubert Soyer;L. Espeholt;Wojciech M. Czarnecki;Simon Schmitt;H. V. Hasselt
中科院分区:
其他
文献类型:
--
作者:
Matteo Hessel;Hubert Soyer;L. Espeholt;Wojciech M. Czarnecki;Simon Schmitt;H. V. Hasselt

文献摘要

被引文献

相似文献

强化学习(RL)社区在设计能够在特定任务上超越人类性能的算法方面取得了长足的进步。这些算法每次只训练一个任务,每个新任务都需要训练一个全新的代理实例。这意味着学习算法是通用的,但每个解都不是;每个代理只能解决它所训练的一个任务。在这项工作中,我们研究了学习一次掌握多个顺序决策任务的问题。多任务学习的一个普遍问题是,必须在竞争单一学习系统有限资源的多个任务的需求之间找到平衡。许多学习算法可能会被要解决的任务集中的某些任务分散注意力。这类任务在学习过程中显得更加突出,例如,因为任务内奖励的密度或大小。这导致算法以牺牲通用性为代价将重点放在那些显著的任务上。我们建议自动调整每个任务对代理更新的贡献,以便所有任务对学习动态具有类似的影响。这导致了在一套57种不同的雅达利游戏中学习玩所有游戏的最先进表现。令人兴奋的是,我们的方法学习了一个单一的训练有素的策略-使用单一的一组权重-超过了人类表现的中位数。据我们所知,在这个多任务领域,这是第一次有单个代理超过人类的表现。同样的方法还在3D强化学习平台DeepMind Lab的一组30个任务上展示了最先进的性能。
The reinforcement learning (RL) community has made great strides in designing algorithms capable of exceeding human performance on specific tasks. These algorithms are mostly trained one task at the time, each new task requiring to train a brand new agent instance. This means the learning algorithm is general, but each solution is not; each agent can only solve the one task it was trained on. In this work, we study the problem of learning to master not one but multiple sequentialdecision tasks at once. A general issue in multi-task learning is that a balance must be found between the needs of multiple tasks competing for the limited resources of a single learning system. Many learning algorithms can get distracted by certain tasks in the set of tasks to solve. Such tasks appear more salient to the learning process, for instance because of the density or magnitude of the in-task rewards. This causes the algorithm to focus on those salient tasks at the expense of generality. We propose to automatically adapt the contribution of each task to the agent’s updates, so that all tasks have a similar impact on the learning dynamics. This resulted in state of the art performance on learning to play all games in a set of 57 diverse Atari games. Excitingly, our method learned a single trained policy - with a single set of weights - that exceeds median human performance. To our knowledge, this was the first time a single agent surpassed human-level performance on this multi-task domain. The same approach also demonstrated state of the art performance on a set of 30 tasks in the 3D reinforcement learning platform DeepMind Lab.