Reinforcement learning for a biped robot based on a CPG-actor-critic method
Reinforcement learning for a biped robot based on a CPG-actor-critic method
复制标题
DOI:
10.1016/j.neunet.2007.01.002
复制
发表时间:
2007-08
期刊:
影响因子:
--
通讯作者:
Yutaka Nakamura;Takeshi Mori;Masa-aki Sato;S. Ishii
中科院分区:
文献类型:
--
作者:
Yutaka Nakamura;Takeshi Mori;Masa-aki Sato;S. Ishii
Animals’ rhythmic movements, such as locomotion, are considered to be controlled by neural circuits called central pattern generators (CPGs), which generate oscillatory signals. Motivated by this biological mechanism, studies have been conducted on the rhythmic movements controlled by CPG. As an autonomous learning framework for a CPG controller, we propose in this article a reinforcement learning method we call the “CPG-actor-critic” method. This method introduces a new architecture to the actor, and its training is roughly based on a stochastic policy gradient algorithm presented recently. We apply this method to an automatic acquisition problem of control for a biped robot. Computer simulations show that training of the CPG can be successfully performed by our method, thus allowing the biped robot to not only walk stably but also adapt to environmental changes.