Constructing continuous action space from basis functions for fast and stable reinforcement learning

Constructing continuous action space from basis functions for fast and stable reinforcement learning
复制标题

从基函数构建连续动作空间以实现快速稳定的强化学习

DOI:
10.1109/roman.2009.5326234
复制
发表时间:
2009
期刊:
RO-MAN 2009 - The 18th IEEE International Symposium on Robot and Human Interactive Communication
影响因子:
--
通讯作者:
T. Ogasawara
T. Ogasawara
中科院分区:
--
文献类型:
--
作者:
Akihiko Yamaguchi;J. Takamatsu;T. Ogasawara

文献摘要

参考文献

被引文献

相似文献

本文提出了一种新的用于线拟合强化学习(RL)的连续动作空间[1]。线拟合具有与基于动作值函数的RL算法一起使用的期望特征。然而,线拟合变得不稳定,导致改变的参数的行动。此外,所获得的行为高度依赖于参数的初始值。提出的动作空间是从山口等人提出的DCOB扩展而来的。[2],其中离散动作集是从给定的基函数生成的。基于DCOB,我们施加一些约束的参数,以获得稳定性。此外,我们还描述了一个适当的方法来初始化的参数。仿真结果表明,该方法优于线拟合。另一方面,所提出的方法的性能是相同的,或劣于DCOB。本文还对这一结果进行了讨论。
This paper presents a new continuous action space for reinforcement learning (RL) with the wire-fitting [1]. The wire-fitting has a desirable feature to be used with action value function based RL algorithms. However, the wire-fitting becomes unstable caused by changing the parameters of actions. Furthermore, the acquired behavior highly depend on the initial values of the parameters. The proposed action space is expanded from the DCOB, proposed by Yamaguchi et al. [2], where the discrete action set is generated from given basis functions. Based on the DCOB, we apply some constraints to the parameters in order to obtain stability. Furthermore, we also describe a proper way to initialize the parameters. The simulation results demonstrate that the proposed method outperforms the wire-fitting. On the other hand, the resulting performance of the proposed method is the same as, or inferior to the DCOB. This paper also discuss about this result.
DOI: 10.1016/j.robot.2003.11.006
发表时间: 2003-09
期刊: Robotics Auton. Syst.
影响因子: --
作者:
T. Kondo;Koji Ito
通讯作者: T. Kondo;Koji Ito