Reinforcement learning of a continuous motor sequence with hidden states

Reinforcement learning of a continuous motor sequence with hidden states
复制标题

DOI:
10.1163/156855307781389365
复制
发表时间:
2007-01
期刊:
影响因子:
2
通讯作者:
H. Arie;T. Ogata;J. Tani;S. Sugano
H. Arie;T. Ogata;J. Tani;S. Sugano
中科院分区:
计算机科学4区
文献类型:
--
作者:
H. Arie;T. Ogata;J. Tani;S. Sugano

文献摘要

被引文献

相似文献

强化学习是一种无监督学习方案,在该方案中,机器人期望通过基于奖励信号的自我探索来获得行为技能。然而,将传统的强化学习算法应用到机器人的运动控制任务中存在一些困难,因为大多数算法涉及离散状态空间,并且基于状态的完全可观察性假设。现实世界的环境通常具有部分可观察性;因此,机器人必须估计不可观察的隐藏状态。本文提出了一种将强化学习算法与连续时间递归神经网络(CTRNN)的学习算法相结合的方法来解决这两个问题。CTRNN可以在连续的时间和空间域中学习时空结构,并通过自组织适当的内部记忆结构来保存上下文流。这使得机器人能够处理隐藏状态问题。我们对不含转速信息的摆锤摆动任务进行了实验。因此,使用所提出的算法在数百次试验中完成了这项任务。此外,本文还证明了将钟摆的转速信息作为一种隐藏状态,在上下文神经元的激活上进行估计和编码。
Reinforcement learning is the scheme for unsupervised learning in which robots are expected to acquire behavior skills through self-explorations based on reward signals. There are some difficulties, however, in applying conventional reinforcement learning algorithms to motion control tasks of a robot because most algorithms are concerned with discrete state space and based on the assumption of complete observability of the state. Real-world environments often have partial observablility; therefore, robots have to estimate the unobservable hidden states. This paper proposes a method to solve these two problems by combining the reinforcement learning algorithm and a learning algorithm for a continuous time recurrent neural network (CTRNN). The CTRNN can learn spatio-temporal structures in a continuous time and space domain, and can preserve the contextual flow by a self-organizing appropriate internal memory structure. This enables the robot to deal with the hidden state problem. We carried out an experiment on the pendulum swing-up task without rotational speed information. As a result, this task is accomplished in several hundred trials using the proposed algorithm. In addition, it is shown that the information about the rotational speed of the pendulum, which is considered as a hidden state, is estimated and encoded on the activation of a context neuron.