Hierarchical Reinforcement Learning via Advantage-Weighted Information Maximization

Hierarchical Reinforcement Learning via Advantage-Weighted Information Maximization
复制标题

DOI:
--
复制
发表时间:
2019-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Takayuki Osa;Voot Tangkaratt;Masashi Sugiyama
Takayuki Osa;Voot Tangkaratt;Masashi Sugiyama
中科院分区:
其他
文献类型:
--
作者:
Takayuki Osa;Voot Tangkaratt;Masashi Sugiyama

文献摘要

被引文献

相似文献

现实世界的任务往往是高度结构化的。层次强化学习(HRL)作为一种利用强化学习(RL)中给定任务的层次结构的方法引起了研究兴趣。然而,识别增强RL性能的分层策略结构并不是一项微不足道的任务。在本文中,我们提出了一个HRL方法,学习一个层次的政策,使用互信息最大化的潜在变量。我们的方法可以被解释为一种学习状态-动作空间的离散和潜在表示的方法。为了学习对应于优势函数模式的期权策略,我们引入了加权重要性抽样。在我们的HRL方法中,门控策略基于期权价值函数学习选择期权策略,并且这些期权策略基于确定性策略梯度方法进行优化。这个框架是通过利用标准RL中的整体策略和HRL中的分层策略之间的类比,通过使用确定性选项策略而得出的。实验结果表明,我们的HRL方法可以学习的选项的多样性,它可以提高RL在连续控制任务的性能。
Real-world tasks are often highly structured. Hierarchical reinforcement learning (HRL) has attracted research interest as an approach for leveraging the hierarchical structure of a given task in reinforcement learning (RL). However, identifying the hierarchical policy structure that enhances the performance of RL is not a trivial task. In this paper, we propose an HRL method that learns a latent variable of a hierarchical policy using mutual information maximization. Our approach can be interpreted as a way to learn a discrete and latent representation of the state-action space. To learn option policies that correspond to modes of the advantage function, we introduce advantage-weighted importance sampling. In our HRL method, the gating policy learns to select option policies based on an option-value function, and these option policies are optimized based on the deterministic policy gradient method. This framework is derived by leveraging the analogy between a monolithic policy in standard RL and a hierarchical policy in HRL by using a deterministic option policy. Experimental results indicate that our HRL approach can learn a diversity of options and that it can enhance the performance of RL in continuous control tasks.