FeUdal Networks for Hierarchical Reinforcement Learning

FeUdal Networks for Hierarchical Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2017-03
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Vezhnevets;Simon Osindero;T. Schaul;N. Heess;Max Jaderberg;David Silver;K. Kavukcuoglu
A. Vezhnevets;Simon Osindero;T. Schaul;N. Heess;Max Jaderberg;David Silver;K. Kavukcuoglu
中科院分区:
其他
文献类型:
--
作者:
A. Vezhnevets;Simon Osindero;T. Schaul;N. Heess;Max Jaderberg;David Silver;K. Kavukcuoglu

文献摘要

被引文献

相似文献

我们介绍了封建网络(FuNs):一种用于分层强化学习的新架构。我们的方法受到Dayan和Hinton的封建强化学习建议的启发,并通过跨多个级别解耦端到端学习来获得强大和有效性-允许它利用不同的时间分辨率。我们的框架使用了一个Manager模块和一个Worker模块。管理器以较低的时间分辨率操作,并设置抽象目标,这些目标由工作器传达并执行。Worker在环境的每一个滴答声中生成原始动作。FuN的解耦结构带来了几个好处——除了促进长期信用分配外,它还鼓励与管理者设定的不同目标相关的子政策的出现。这些属性允许FuN在涉及长期信用分配或记忆的任务上显著优于强大的基线代理。我们在ATARI套件和3D DeepMind实验室环境中演示了我们提出的系统在一系列任务中的性能。
We introduce FeUdal Networks (FuNs): a novel architecture for hierarchical reinforcement learning. Our approach is inspired by the feudal reinforcement learning proposal of Dayan and Hinton, and gains power and efficacy by decoupling end-to-end learning across multiple levels -- allowing it to utilise different resolutions of time. Our framework employs a Manager module and a Worker module. The Manager operates at a lower temporal resolution and sets abstract goals which are conveyed to and enacted by the Worker. The Worker generates primitive actions at every tick of the environment. The decoupled structure of FuN conveys several benefits -- in addition to facilitating very long timescale credit assignment it also encourages the emergence of sub-policies associated with different goals set by the Manager. These properties allow FuN to dramatically outperform a strong baseline agent on tasks that involve long-term credit assignment or memorisation. We demonstrate the performance of our proposed system on a range of tasks from the ATARI suite and also from a 3D DeepMind Lab environment.