Developing reinforcement learning for adaptive co-construction of continuous high-dimensional state and action spaces

Developing reinforcement learning for adaptive co-construction of continuous high-dimensional state and action spaces
复制标题

开发强化学习以自适应共建连续高维状态和动作空间

DOI:
10.1007/s10015-012-0041-5
复制
发表时间:
2012
影响因子:
0.9
通讯作者:
H. Tamaki
H. Tamaki
中科院分区:
--
文献类型:
--
作者:
M. Nagayoshi;H. Murao;H. Tamaki

文献摘要

参考文献

被引文献

相似文献

工程师和研究人员越来越关注强化学习(RL)作为实现自适应和自治分散系统的关键技术。然而,一般来说,将RL投入实际使用并不容易。我们的方法主要处理的问题,设计状态和动作空间。以前,在其他空间已经固定之后,已经提出了被称为“状态空间滤波器”的自适应状态空间构造方法和被称为“切换RL”的自适应动作空间构造方法。然后,我们已经重组这两种构造方法作为一种方法,通过处理前一种方法和后一种方法作为一种组合的方法,用于模仿婴儿的知觉和运动的发展,我们提出了一种方法,这是基于引入和参考“熵”。在本文中,进行了计算实验,使用所谓的“机器人导航问题”与三维连续的状态空间和二维连续的动作空间,这是更复杂的比所谓的“路径规划问题”。结果表明,该方法是有效的.
Engineers and researchers are paying more attention to reinforcement learning (RL) as a key technique for realizing adaptive and autonomous decentralized systems. In general, however, it is not easy to put RL into practical use. Our approach mainly deals with the problem of designing state and action spaces. Previously, an adaptive state space construction method which is called a “state space filter” and an adaptive action space construction method which is called “switching RL”, have been proposed after the other space has been fixed. Then, we have reconstituted these two construction methods as one method by treating the former method and the latter method as a combined method for mimicking an infant’s perceptual and motor developments and we have proposed a method which is based on introducing and referring to “entropy”. In this paper, a computational experiment was conducted using a so-called “robot navigation problem” with three-dimensional continuous state space and two-dimensional continuous action space which is more complicated than a so-called “path planning problem”. As a result, the validity of the proposed method has been confirmed.
强化学习中状态和行动空间的自适应共建
DOI: --
发表时间: 2011
期刊: Proc. of the 16^<th> Int. Symp. on Artificial Life and Robotics
影响因子: --
作者:
Masato Nagayoshi;Hajime Murao;Hisashi Tamaki
通讯作者: Hisashi Tamaki