Learning CPG-based biped locomotion with a policy gradient method: Application to a humanoid robot

Learning CPG-based biped locomotion with a policy gradient method: Application to a humanoid robot
复制标题

DOI:
10.1177/0278364907084980
复制
发表时间:
2008-02-01
影响因子:
9.2
通讯作者:
Cheng, Gordon
Cheng, Gordon
中科院分区:
计算机科学2区
文献类型:
--
作者:
Endo, Gen;Morimoto, Jun;Cheng, Gordon

文献摘要

被引文献

相似文献

在本文中,我们描述了一个学习框架的中央模式发生器(CPG)为基础的运动控制器使用的政策梯度方法。我们在这项研究中的目标是实现基于CPG的三维硬件人形机器人的步行,并开发一个有效的学习算法与CPG通过减少用于学习的状态空间的维数。我们证明了一个适当的反馈控制器可以获得在几千次试验的数值模拟和数值模拟中获得的控制器实现了稳定的步行与物理机器人在真实的世界。数值模拟和硬件实验评估步行速度和稳定性。结果表明,学习算法能够适应环境的变化。此外,我们提出了一个在线学习计划的初始政策的硬件机器人,以提高控制器在200次迭代。
In this paper we describe a learning framework for a central pattern generator (CPG)-based biped locomotion controller using a policy gradient method. Our goals in this study are to achieve CPG-based biped walking with a 3D hardware humanoid and to develop an efficient learning algorithm with CPG by reducing the dimensionality of the state space used for learning. We demonstrate that an appropriate feedback controller can be acquired within a few thousand trials by numerical simulations and the controller obtained in numerical simulation achieves stable walking with a physical robot in the real world. Numerical simulations and hardware experiments evaluate the walking velocity and stability. The results suggest that the learning algorithm is capable of adapting to environmental changes. Furthermore, we present an online learning scheme with an initial policy for a hardware robot to improve the controller within 200 iterations.