Efficient Reinforcement Learning with Hierarchies of Machines by Leveraging Internal Transitions

Efficient Reinforcement Learning with Hierarchies of Machines by Leveraging Internal Transitions
复制标题

DOI:
10.24963/ijcai.2017/196
复制
发表时间:
2017-08
期刊:
--
影响因子:
--
通讯作者:
Aijun Bai;Stuart J. Russell
Aijun Bai;Stuart J. Russell
中科院分区:
其他
文献类型:
--
作者:
Aijun Bai;Stuart J. Russell

文献摘要

相似文献

在层次强化学习的背景下,抽象机层次(HAM)的思想是将部分策略编写为一组具有未指定选择状态的层次有限状态机,并使用强化学习来学习此部分策略的最佳完成。给定具有深层层次结构的HAM,通常存在许多内部转换,其中一台机器在环境状态不变的情况下调用另一台机器。在本文中,我们提出了一种新的分层强化学习算法,自动发现这样的内部转换,并在Q值的计算中递归地缩短它们。由此产生的HAMQ-INT算法在基准Taxi域和更复杂的RoboCup Keepaway域上的性能明显优于现有技术。
In the context of hierarchical reinforcement learning, the idea of hierarchies of abstract machines (HAMs) is to write a partial policy as a set of hierarchical finite state machines with unspecified choice states, and use reinforcement learning to learn an optimal completion of this partial policy. Given a HAM with deep hierarchical structure, there often exist many internal transitions where a machine calls another machine with the environment state unchanged. In this paper, we propose a new hierarchical reinforcement learning algorithm that automatically discovers such internal transitions, and shortcircuits them recursively in the computation of Q values. The resulting HAMQ-INT algorithm outperforms the state of the art significantly on the benchmark Taxi domain and a much more complex RoboCup Keepaway domain.