Monte carlo bayesian hierarchical reinforcement learning

Monte carlo bayesian hierarchical reinforcement learning
复制标题

蒙特卡洛贝叶斯分层强化学习

DOI:
--
复制
发表时间:
2014
期刊:
Adaptive Agents and Multi-Agent Systems
影响因子:
--
通讯作者:
W. Ertel
W. Ertel
中科院分区:
--
文献类型:
--
作者:
Ngo Anh Vien;H. Ngo;W. Ertel

文献摘要

被引文献

相似文献

在本文中,我们提出使用层次动作分解使基于贝叶斯模型的强化学习在实践中更加高效和可行。我们将贝叶斯分层强化学习表述为部分可观察的半马尔可夫决策过程(POSMDP)。主POSMDP任务被划分为POSMDP子任务的层次结构;较低层次的子任务首先得到解决,然后才是更高层次的子任务。我们从一个先验信念中采样,为每个POSMDP建立一个近似模型,然后使用蒙特卡罗值迭代和宏动作求解器进行求解。实验结果表明,我们的算法在奖励方面,特别是在求解时间方面,都明显优于平面BRL算法,至少好一个数量级。
In this paper, we propose to use hierarchical action decomposition to make Bayesian model-based reinforcement learning more efficient and feasible in practice. We formulate Bayesian hierarchical reinforcement learning as a partially observable semi-Markov decision process (POSMDP). The main POSMDP task is partitioned into a hierarchy of POSMDP subtasks; lower-level subtasks get solved first, then higher-level ones. We sample from a prior belief to build an approximate model for each POSMDP, then solve using Monte Carlo Value Iteration with Macro-Actions solver. Experimental results show that our algorithm performs significantly better than that of flat BRL in terms of both reward, and especially solving time, in at least one order of magnitude.