Abstract Value Iteration for Hierarchical Reinforcement Learning

Abstract Value Iteration for Hierarchical Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
--
影响因子:
--
通讯作者:
Kishor Jothimurugan;O. Bastani;R. Alur
Kishor Jothimurugan;O. Bastani;R. Alur
中科院分区:
其他
文献类型:
--
作者:
Kishor Jothimurugan;O. Bastani;R. Alur

文献摘要

被引文献

相似文献

我们提出了一种新的分层强化学习框架,用于连续状态和动作空间的控制。在我们的框架中,用户指定的子目标区域的状态的子集,然后,我们(i)学习选项,作为这些子目标区域之间的过渡,以及(ii)在由此产生的抽象决策过程(ADP)构建一个高层次的计划。一个关键的挑战是,ADP可能不是马尔可夫,我们提出了两个算法,规划在ADP解决。我们的第一个算法是保守的,使我们能够证明其性能的理论保证,这有助于通知子目标区域的设计。我们的第二个算法是一个实用的交织在抽象层面的规划和学习在具体的水平。在我们的实验中,我们证明了我们的方法在几个具有挑战性的基准测试中优于最先进的分层强化学习算法。
We propose a novel hierarchical reinforcement learning framework for control with continuous state and action spaces. In our framework, the user specifies subgoal regions which are subsets of states; then, we (i) learn options that serve as transitions between these subgoal regions, and (ii) construct a high-level plan in the resulting abstract decision process (ADP). A key challenge is that the ADP may not be Markov, which we address by proposing two algorithms for planning in the ADP. Our first algorithm is conservative, allowing us to prove theoretical guarantees on its performance, which help inform the design of subgoal regions. Our second algorithm is a practical one that interweaves planning at the abstract level and learning at the concrete level. In our experiments, we demonstrate that our approach outperforms state-of-the-art hierarchical reinforcement learning algorithms on several challenging benchmarks.