Data-Efficient and Safe Learning for Humanoid Locomotion Aided by a Dynamic Balancing Model

Data-Efficient and Safe Learning for Humanoid Locomotion Aided by a Dynamic Balancing Model
复制标题

DOI:
10.1109/lra.2020.2990743
复制
发表时间:
2019-06
影响因子:
5.2
通讯作者:
Junhyeok Ahn;Jaemin Lee;L. Sentis
Junhyeok Ahn;Jaemin Lee;L. Sentis
中科院分区:
计算机科学2区
文献类型:
--
作者:
Junhyeok Ahn;Jaemin Lee;L. Sentis

文献摘要

被引文献

相似文献

在这封信中,我们提出了一种新的马尔可夫决策过程(MDP),用于在动态平衡模型的辅助下安全和数据高效地学习类人运动。在我们以前对两足动物运动的研究中,我们依赖于一个低维的机器人模型,该模型通常用于高级行走模式生成器(WPG)。然而,由于全阶模型和简化模型之间的差异,低级反馈控制器不能精确跟踪期望的足迹位置。在这项研究中,我们建议通过用强化学习来补充WPG来缓解这个问题。更具体地说,我们提出了一种由WPG、神经网络和安全控制器组成的结构化足迹控制方法。WPG提供了一种分析方法,促进了有效学习,同时神经网络最大化了长期回报,安全控制器基于步骤捕获和控制屏障功能的使用鼓励安全探索。我们的贡献包括:(1)用于运动的结构化学习控制方法;(2)使用基于物理模型的数据高效和安全的学习过程来改善步行;(3)该过程可扩展到各种类型的类人机器人和步行。
In this letter, we formulate a novel Markov Decision Process (MDP) for safe and data-efficient learning for humanoid locomotion aided by a dynamic balancing model. In our previous studies of biped locomotion, we relied on a low-dimensional robot model, commonly used in high-level Walking Pattern Generators (WPGs). However, a low-level feedback controller cannot precisely track desired footstep locations due to the discrepancies between the full order model and the simplified model. In this study, we propose mitigating this problem by complementing a WPG with reinforcement learning. More specifically, we propose a structured footstep control method consisting of a WPG, a neural network, and a safety controller. The WPG provides an analytical method that promotes efficient learning while the neural network maximizes long-term rewards, and the safety controller encourages safe exploration based on step capturability and the use of control-barrier functions. Our contributions include the following (1) a structured learning control method for locomotion, (2) a data-efficient and safe learning process to improve walking using a physics-based model, and (3) the scalability of the procedure to various types of humanoid robots and walking.