Reinforcement learning for a biped robot based on a CPG-actor-critic method

Reinforcement learning for a biped robot based on a CPG-actor-critic method
复制标题

DOI:
10.1016/j.neunet.2007.01.002
复制
发表时间:
2007-08
期刊:
Neural networks : the official journal of the International Neural Network Society
影响因子:
--
通讯作者:
Yutaka Nakamura;Takeshi Mori;Masa-aki Sato;S. Ishii
Yutaka Nakamura;Takeshi Mori;Masa-aki Sato;S. Ishii
中科院分区:
其他
文献类型:
--
作者:
Yutaka Nakamura;Takeshi Mori;Masa-aki Sato;S. Ishii

文献摘要

被引文献

相似文献

动物的有节奏的运动,如运动,被认为是由称为中央模式发生器(CPG)的神经回路控制的,它产生振荡信号。受这种生物学机制的启发,人们对CPG控制的节律性运动进行了研究。作为一个自主学习框架的CPG控制器,我们在这篇文章中提出了一种强化学习方法,我们称之为“CPG-actor-critic”方法。该方法引入了一种新的结构的演员,其训练大致是基于最近提出的随机策略梯度算法。我们将此方法应用于一个机器人控制的自动获取问题。计算机模拟表明,我们的方法可以成功地进行训练的CPG,从而使机器人不仅行走稳定,而且还能适应环境的变化。
Animals’ rhythmic movements, such as locomotion, are considered to be controlled by neural circuits called central pattern generators (CPGs), which generate oscillatory signals. Motivated by this biological mechanism, studies have been conducted on the rhythmic movements controlled by CPG. As an autonomous learning framework for a CPG controller, we propose in this article a reinforcement learning method we call the “CPG-actor-critic” method. This method introduces a new architecture to the actor, and its training is roughly based on a stochastic policy gradient algorithm presented recently. We apply this method to an automatic acquisition problem of control for a biped robot. Computer simulations show that training of the CPG can be successfully performed by our method, thus allowing the biped robot to not only walk stably but also adapt to environmental changes.