Learning Multi-Level Hierarchies with Hindsight

Learning Multi-Level Hierarchies with Hindsight
复制标题

DOI:
--
复制
发表时间:
2017-12
期刊:
--
影响因子:
--
通讯作者:
Andrew Levy;G. Konidaris;Robert W. Platt;Kate Saenko
Andrew Levy;G. Konidaris;Robert W. Platt;Kate Saenko
中科院分区:
其他
文献类型:
--
作者:
Andrew Levy;G. Konidaris;Robert W. Platt;Kate Saenko

文献摘要

相似文献

分级代理比非分级代理具有以更高的样本效率解决顺序决策任务的潜力,因为分级代理可以将任务分解成只需要较短决策序列的子任务集。为了实现更快学习的潜力,分层代理需要能够并行学习它们的多个级别的策略,以便可以同时解决这些更简单的子问题。然而,并行学习多个级别的策略是困难的,因为它本质上是不稳定的:层次结构中一个级别的策略的变化可能会导致层次结构中较高级别的过渡和奖励功能发生变化,从而难以联合学习多个级别的策略。本文提出了一种新的层次化强化学习(HRL)框架--层次化参与者-批评者(HAC),它可以克服智能体试图联合学习多层策略时出现的不稳定性问题。HAC背后的主要思想是通过训练每个级别来独立于较低级别来训练层次的每个级别,就好像较低级别的策略已经是最优的一样。我们在网格世界和模拟机器人领域的实验表明,相对于其他非层次化和层次化的方法,我们的方法可以显著地加速学习。事实上,我们的框架是第一个成功地在具有连续状态和动作空间的任务中并行学习3级层次结构。
Hierarchical agents have the potential to solve sequential decision making tasks with greater sample efficiency than their non-hierarchical counterparts because hierarchical agents can break down tasks into sets of subtasks that only require short sequences of decisions. In order to realize this potential of faster learning, hierarchical agents need to be able to learn their multiple levels of policies in parallel so these simpler subproblems can be solved simultaneously. Yet, learning multiple levels of policies in parallel is hard because it is inherently unstable: changes in a policy at one level of the hierarchy may cause changes in the transition and reward functions at higher levels in the hierarchy, making it difficult to jointly learn multiple levels of policies. In this paper, we introduce a new Hierarchical Reinforcement Learning (HRL) framework, Hierarchical Actor-Critic (HAC), that can overcome the instability issues that arise when agents try to jointly learn multiple levels of policies. The main idea behind HAC is to train each level of the hierarchy independently of the lower levels by training each level as if the lower level policies are already optimal. We demonstrate experimentally in both grid world and simulated robotics domains that our approach can significantly accelerate learning relative to other non-hierarchical and hierarchical methods. Indeed, our framework is the first to successfully learn 3-level hierarchies in parallel in tasks with continuous state and action spaces.