Learning CPG-based biped locomotion with a policy gradient method

Learning CPG-based biped locomotion with a policy gradient method
复制标题

DOI:
10.1016/j.robot.2006.05.012
复制
发表时间:
2005-12
期刊:
5th IEEE-RAS International Conference on Humanoid Robots, 2005.
影响因子:
--
通讯作者:
Takamitsu Matsubara;Jun Morimoto;Jun Nakanishi;Masa-aki Sato;Kenji Doya
Takamitsu Matsubara;Jun Morimoto;Jun Nakanishi;Masa-aki Sato;Kenji Doya
中科院分区:
其他
文献类型:
--
作者:
Takamitsu Matsubara;Jun Morimoto;Jun Nakanishi;Masa-aki Sato;Kenji Doya

文献摘要

相似文献

在本文中,我们提出了一个学习框架,基于CPG的策略梯度方法的移动。我们证明,适当的感官反馈,以调整的CPG(中央模式发生器)的节奏,可以学习使用所提出的方法在几百次试验中的模拟。我们调查的线性稳定性的周期轨道的收购步行模式考虑其近似返回地图。此外,我们将数值模拟中获得的控制器应用于我们的物理5连杆机器人,以经验评估在真实的环境中行走的鲁棒性。实验结果表明,机器人能够成功地走使用所获得的控制器,即使在环境变化的情况下,通过放置一个跷跷板状的金属片在地面上和一个参数变化的机器人动力学与一个额外的重量在小腿上,这是没有建模的数值模拟。
In this paper, we propose a learning framework for CPG-based biped locomotion with a policy gradient method. We demonstrate that appropriate sensory feedback to adjust the rhythm of the CPG (Central Pattern Generator) can be learned using the proposed method within a few hundred trials in simulations. We investigate linear stability of a periodic orbit of the acquired walking pattern considering its approximated return map. Furthermore, we apply the controllers acquired in numerical simulations to our physical 5-link biped robot in order to empirically evaluate the robustness of walking in the real environment. Experimental results demonstrate that the robot was able to successfully walk using the acquired controllers even in the cases of an environmental change by placing a seesaw-like metal sheet on the ground and a parametric change of the robot dynamics with an additional weight on a shank, which was not modeled in the numerical simulations.