Kernel-based direct policy search reinforcement learning based on variational Bayesian inference

Kernel-based direct policy search reinforcement learning based on variational Bayesian inference
复制标题

基于变分贝叶斯推理的核直接策略搜索强化学习

DOI:
10.1109/candarw.2019.00040
复制
发表时间:
2019
期刊:
Proceedings of 2019 Seventh International Symposium on Computing and Networking Workshops (CANDARW)
影响因子:
--
通讯作者:
Hiroshi Okumura
Hiroshi Okumura
中科院分区:
--
文献类型:
--
作者:
Nobuhiko Yamaguchi;Osamu Fukuda;Hiroshi Okumura

文献摘要

相似文献

直接策略搜索是一种很有前途的强化学习框架,特别是用于控制连续的高维系统。作为直接策略搜索的一种,提出了基于变分贝叶斯推理的直接策略搜索强化学习(VBRL)。 VBRL算法基于变分贝叶斯推理来估计策略参数,因此避免了过拟合问题。在本文中,我们提出了使用核方法技术对 VBRL 模型的扩展,我们称之为 K-VBRL。所提出的 K-VBRL 的性能通过两个山地车任务实验进行评估。这些实验强调 K-VBRL 产生更高的平均回报并且优于传统的 VBRL。
Direct policy search is a promising reinforcement learning framework in particular for controlling continuous, high-dimensional systems. As one of direct policy search, direct policy search reinforcement learning based on variational Bayesian inference (VBRL) was proposed. The VBRL algorithm estimates the policy parameter based on variational Bayesian inference and is therefore avoid overfitting problem. In this paper, we propose an extension of the VBRL model using techniques of kernel methods, which we call K-VBRL. The performance of the proposed K-VBRL is assessed in two experiments with mountain car task. These experiments highlight the K-VBRL produces higher average return and outperforms the conventional VBRL.