Kernel-based direct policy search reinforcement learning based on variational Bayesian inference
Kernel-based direct policy search reinforcement learning based on variational Bayesian inference
复制标题
基于变分贝叶斯推理的核直接策略搜索强化学习
DOI:
10.1109/candarw.2019.00040
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Hiroshi Okumura
中科院分区:
文献类型:
--
作者:
Nobuhiko Yamaguchi;Osamu Fukuda;Hiroshi Okumura
Direct policy search is a promising reinforcement learning framework in particular for controlling continuous, high-dimensional systems. As one of direct policy search, direct policy search reinforcement learning based on variational Bayesian inference (VBRL) was proposed. The VBRL algorithm estimates the policy parameter based on variational Bayesian inference and is therefore avoid overfitting problem. In this paper, we propose an extension of the VBRL model using techniques of kernel methods, which we call K-VBRL. The performance of the proposed K-VBRL is assessed in two experiments with mountain car task. These experiments highlight the K-VBRL produces higher average return and outperforms the conventional VBRL.