Efficient Reinforcement Learning with Hierarchies of Machines by Leveraging Internal Transitions
Efficient Reinforcement Learning with Hierarchies of Machines by Leveraging Internal Transitions
复制标题
DOI:
10.24963/ijcai.2017/196
复制
发表时间:
2017-08
期刊:
影响因子:
--
通讯作者:
Aijun Bai;Stuart J. Russell
中科院分区:
文献类型:
--
作者:
Aijun Bai;Stuart J. Russell
In the context of hierarchical reinforcement learning, the idea of hierarchies of abstract machines (HAMs) is to write a partial policy as a set of hierarchical finite state machines with unspecified choice states, and use reinforcement learning to learn an optimal completion of this partial policy. Given a HAM with deep hierarchical structure, there often exist many internal transitions where a machine calls another machine with the environment state unchanged. In this paper, we propose a new hierarchical reinforcement learning algorithm that automatically discovers such internal transitions, and shortcircuits them recursively in the computation of Q values. The resulting HAMQ-INT algorithm outperforms the state of the art significantly on the benchmark Taxi domain and a much more complex RoboCup Keepaway domain.