Probabilistic inference for determining options in reinforcement learning

Probabilistic inference for determining options in reinforcement learning
复制标题

DOI:
10.1007/s10994-016-5580-x
复制
发表时间:
2016-08
期刊:
影响因子:
7.5
通讯作者:
Christian Daniel;H. V. Hoof;Jan Peters;G. Neumann
Christian Daniel;H. V. Hoof;Jan Peters;G. Neumann
中科院分区:
计算机科学3区
文献类型:
--
作者:
Christian Daniel;H. V. Hoof;Jan Peters;G. Neumann

文献摘要

被引文献

相似文献

需要许多顺序决策或复杂解决方案的任务很难使用传统的强化学习算法来解决。基于半马尔可夫决策过程设置(SMDP)和选项框架,我们提出了一个模型,旨在减轻这些问题。代理不是学习单个整体策略,而是学习一组更简单的子策略以及每个子策略的启动和终止概率。虽然现有的选项学习算法经常需要手动规范的组件,如子政策,我们提出了一个算法,推断所有相关组件的选项框架数据。此外,所提出的方法是基于参数化的选项表示,并与当前的政策搜索方法,这是特别适合于连续的现实世界的任务相结合。我们目前的结果SMDP离散以及连续的状态动作空间。实验结果表明,该算法能够将简单的子策略联合收割机组合起来解决复杂任务,并能提高简单任务的学习性能。
Tasks that require many sequential decisions or complex solutions are hard to solve using conventional reinforcement learning algorithms. Based on the semi Markov decision process setting (SMDP) and the option framework, we propose a model which aims to alleviate these concerns. Instead of learning a single monolithic policy, the agent learns a set of simpler sub-policies as well as the initiation and termination probabilities for each of those sub-policies. While existing option learning algorithms frequently require manual specification of components such as the sub-policies, we present an algorithm which infers all relevant components of the option framework from data. Furthermore, the proposed approach is based on parametric option representations and works well in combination with current policy search methods, which are particularly well suited for continuous real-world tasks. We present results on SMDPs with discrete as well as continuous state-action spaces. The results show that the presented algorithm can combine simple sub-policies to solve complex tasks and can improve learning performance on simpler tasks.