Regularizing Reinforcement Learning with State Abstraction

Regularizing Reinforcement Learning with State Abstraction
复制标题

DOI:
10.1109/iros.2018.8594201
复制
发表时间:
2018-10
期刊:
2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
R. Akrour;Filipe Veiga;Jan Peters;G. Neumann
R. Akrour;Filipe Veiga;Jan Peters;G. Neumann
中科院分区:
其他
文献类型:
--
作者:
R. Akrour;Filipe Veiga;Jan Peters;G. Neumann

文献摘要

相似文献

离散强化学习设置中的状态抽象将共享类似最佳动作的状态聚类,以产生更容易解决的决策过程。在本文中,我们将状态抽象的概念推广到连续动作强化学习,将抽象状态定义为一个状态簇,在这个状态簇上存在一个简单形状的近似最优策略。我们提出了一种分层强化学习算法,能够同时找到状态空间聚类和最优子策略在每个集群。所提出的框架的主要优点是通过控制学习策略的行为复杂性来提供一种简单的正则化强化学习的方法。我们将我们的算法应用于几个基准任务和一个机器人触觉操作任务,并表明我们可以通过组合少量线性策略来匹配最先进的深度强化学习性能。
State abstraction in a discrete reinforcement learning setting clusters states sharing a similar optimal action to yield an easier to solve decision process. In this paper, we generalize the concept of state abstraction to continuous action reinforcement learning by defining an abstract state as a state cluster over which a near-optimal policy of simple shape exists. We propose a hierarchical reinforcement learning algorithm that is able to simultaneously find the state space clustering and the optimal sub-policies in each cluster. The main advantage of the proposed framework is to provide a straightforward way of regularizing reinforcement learning by controlling the behavioral complexity of the learned policy. We apply our algorithm on several benchmark tasks and a robot tactile manipulation task and show that we can match state-of-the-art deep reinforcement learning performance by combining a small number of linear policies.