Constructing action set from basis functions for reinforcement learning of robot control

Constructing action set from basis functions for reinforcement learning of robot control
复制标题

从基函数构建机器人控制强化学习的动作集

DOI:
10.1109/robot.2009.5152840
复制
发表时间:
2009
期刊:
2009 IEEE International Conference on Robotics and Automation
影响因子:
--
通讯作者:
T. Ogasawara
T. Ogasawara
中科院分区:
--
文献类型:
--
作者:
Akihiko Yamaguchi;J. Takamatsu;T. Ogasawara

文献摘要

参考文献

被引文献

相似文献

由于控制输入是连续的,连续动作集被用于机器人控制的许多强化学习 (RL) 应用中。然而,离散动作集还具有易于实现以及与一些复杂的 RL 方法(例如 Dyna [1])兼容的优点。然而,问题之一是缺乏在高维输入空间中设计用于机器人控制的离散动作集的一般原则。在本文中,我们建议在给定一组基函数(BF)的情况下构造一个离散动作集。我们设计了动作集,使得动作集的大小与 BF 的数量成正比。该方法可以利用函数逼近器的性质,即在实际的强化学习应用中,BF 的数量不会随着状态空间的维数呈指数增长(例如[2])。因此,建议的动作集的大小不会随着输入空间的维度呈指数增长。我们将具有建议动作集的强化学习应用于机器人导航任务以及爬行和跳跃任务。仿真结果表明,与传统的离散动作集相比,所提出的动作集具有提高学习速度和更好的获得性能的能力的优点。
Continuous action sets are used in many reinforcement learning (RL) applications in robot control since the control input is continuous. However, discrete action sets also have the advantages of ease of implementation and compatibility with some sophisticated RL methods, such as the Dyna [1]. However, one of the problem is the absence of general principles on designing a discrete action set for robot control in higher dimensional input space. In this paper, we propose to construct a discrete action set given a set of basis functions (BFs). We designed the action set so that the size of the set is proportional to the number of the BFs. This method can exploit the function approximator's nature, that is, in practical RL applications, the number of BFs does not increase exponentially with the dimension of the state space (e.g. [2]). Thus, the size of the proposed action set does not increase exponentially with the dimension of the input space. We apply an RL with the proposed action set to a robot navigation task and a crawling and a jumping tasks. The simulation results demonstrate that the proposed action set has the advantages of improved learning speed, and better ability to acquire performance, compared to a conventional discrete action set.
DOI: 10.1016/j.robot.2003.11.006
发表时间: 2003-09
期刊: Robotics Auton. Syst.
影响因子: --
作者:
T. Kondo;Koji Ito
通讯作者: T. Kondo;Koji Ito