Biped Walk Learning Through Playback and Corrective Demonstration

Biped Walk Learning Through Playback and Corrective Demonstration
复制标题

通过回放和纠正演示进行两足步行学习

DOI:
10.1609/aaai.v24i1.7730
复制
发表时间:
2010
期刊:
Proceedings of the AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
M. Veloso
M. Veloso
中科院分区:
--
文献类型:
--
作者:
Çetin Meriçli;M. Veloso

文献摘要

被引文献

相似文献

开发一个强大的,灵活的,闭环的人形机器人步行算法是一项具有挑战性的任务,由于一般的仿人步行的复杂动力学。通常的分析方法是使用物理现实的简化模型。这种方法是部分成功的,因为它们导致机器人行走在不可避免的福尔斯方面的失败。而不是进一步完善的分析模型,在这项工作中,我们调查使用人类纠正示范,因为我们意识到,一个人可以直观地检测到机器人可能会下降。我们贡献了一个两阶段的步行学习方法,我们实验的毕宿五NAO人形机器人。在第一阶段中,机器人步行以下的分析简化步行算法,这是一个黑盒子,我们确定并保存一个步行周期作为关节运动命令。然后,我们展示了机器人如何可以重复和成功地回放记录的运动周期,即使在开环。在第二阶段,我们通过修改记录的步行周期来响应传感器数据来创建闭环步行。该算法学习关节运动校正的开环步行的基础上提供的纠正反馈的人,并在传感器数据,而自主行走。在我们的实验结果中,我们表明,学习的闭环行走策略优于手动调整的闭环策略和开环回放行走,在机器人不跌倒的情况下行驶的距离。
Developing a robust, flexible, closed-loop walking algorithm for a humanoid robot is a challenging task due to the complex dynamics of the general biped walk. Common analytical approaches to biped walk use simplified models of the physical reality. Such approaches are partially successful as they lead to failures of the robot walk in terms of unavoidable falls. Instead of further refining the analytical models, in this work we investigate the use of human corrective demonstrations, as we realize that a human can visually detect when the robot may be falling. We contribute a two-phase biped walk learning approach, which we experiment on the Aldebaran NAO humanoid robot. In the first phase, the robot walks following an analytical simplified walk algorithm, which is used as a black box, and we identify and save a walk cycle as joint motion commands. We then show how the robot can repeatedly and successfully play back the recorded motion cycle, even if in open-loop. In the second phase, we create a closed-loop walk by modifying the recorded walk cycle to respond to sensory data. The algorithm learns joint movement corrections to the open-loop walk based on the corrective feedback provided by a human, and on the sensory data, while walking autonomously. In our experimental results, we show that the learned closed-loop walking policy outperforms a hand-tuned closed-loop policy and the open-loop playback walk, in terms of the distance traveled by the robot without falling.