Hierarchical reinforcement learning with the MAXQ value function decomposition

Hierarchical reinforcement learning with the MAXQ value function decomposition
复制标题

DOI:
10.1613/jair.639
复制
发表时间:
2000-01-01
影响因子:
5
通讯作者:
Dietterich, TG
Dietterich, TG
中科院分区:
计算机科学3区
文献类型:
--
作者:
Dietterich, TG

文献摘要

被引文献

相似文献

本文提出了一种新的分层强化学习方法,该方法基于将目标马尔可夫决策过程(MDP)分解为较小的 MDP 的层次结构,并将目标 MDP 的价值函数分解为较小的 MDP 的价值函数的加法组合。这种分解称为 MAXQ 分解,既具有过程语义(作为子例程层次结构),又具有声明性语义(作为层次策略的值函数的表示)。 MAXQ 统一并扩展了 Singh、Kaelbling、Dayan 和 Hinton 之前关于分层强化学习的工作。它基于这样的假设:程序员可以识别有用的子目标并定义实现这些子目标的子任务。通过定义此类子目标,程序员可以限制强化学习期间需要考虑的策略集。 MAXQ价值函数分解可以表示与给定层次结构一致的任何策略的价值函数。分解还创造了利用状态抽象的机会,以便层次结构中的各个 MDP 可以忽略状态空间的大部分。这对于该方法的实际应用很重要。本文定义了 MAXQ 层次结构,证明了其表征能力的形式化结果,并建立了安全使用状态抽象的五个条件。该论文提出了一种在线无模型学习算法 MAXQ-Q,并证明即使存在五种状态抽象,它也能以概率 1 收敛到一种称为递归最优策略的局部最优策略。该论文通过三个领域的一系列实验评估了 MAXQ 表示和 MAXQ-Q,并通过实验证明 MAXQ-Q(具有状态抽象)比 Q 学习更快地收敛到递归最优策略。事实上,MAXQ 学习值函数的表示有一个重要的好处:它可以通过类似于策略迭代的策略改进步骤的过程来计算和执行改进的非分层策略。论文通过实验证明了这种非分层执行的有效性。最后,本文对相关工作进行了比较,并讨论了分层强化学习中的设计权衡。
This paper presents a new approach to hierarchical reinforcement learning based on decomposing the target Markov decision process (MDP) into a hierarchy of smaller MDPs and decomposing the value function of the target MDP into an additive combination of the value functions of the smaller MDPs. The decomposition, known as the MAXQ decomposition, has both a procedural semantics-as a subroutine hierarchy-and a declarative semantics-as a representation of the value function of a hierarchical policy. MAXQ unifies and extends previous work on hierarchical reinforcement learning by Singh, Kaelbling, and Dayan and Hinton. It is based on the assumption that the programmer can identify useful subgoals and define subtasks that achieve these subgoals. By defining such subgoals, the programmer constrains the set of policies that need to be considered during reinforcement learning. The MAXQ value function decomposition can represent the value function of any policy that is consistent with the given hierarchy. The decomposition also creates opportunities to exploit state abstractions, so that individual MDPs within the hierarchy can ignore large parts of the state space. This is important for the practical application of the method. This paper defines the MAXQ hierarchy, proves formal results on its representational power, and establishes five conditions for the safe use of state abstractions. The paper presents an online model-free learning algorithm, MAXQ-Q, and proves that it converges with probability 1 to a kind of locally-optimal policy known as a recursively optimal policy, even in the presence of the five kinds of state abstraction. The paper evaluates the MAXQ representation and MAXQ-Q through a series of experiments in three domains and shows experimentally that MAXQ-Q (with state abstractions) converges to a recursively optimal policy much faster than at Q learning. The fact that MAXQ learns a representation of the value function has an important benefit: it makes it possible to compute and execute an improved, non-hierarchical policy via a procedure similar to the policy improvement step of policy iteration. The paper demonstrates the effectiveness of this non-hierarchical execution experimentally. Finally, the paper concludes with a comparison to related work and a discussion of the design tradeoffs in hierarchical reinforcement learning.