Reinforcement Learning-Based Cascade Motion Policy Design for Robust 3D Bipedal Locomotion

Reinforcement Learning-Based Cascade Motion Policy Design for Robust 3D Bipedal Locomotion
复制标题

DOI:
10.1109/access.2022.3151771
复制
发表时间:
2022
期刊:
影响因子:
3.9
通讯作者:
Guillermo A. Castillo;Bowen Weng;Wei Zhang;Ayonga Hereid
Guillermo A. Castillo;Bowen Weng;Wei Zhang;Ayonga Hereid
中科院分区:
计算机科学3区
文献类型:
--
作者:
Guillermo A. Castillo;Bowen Weng;Wei Zhang;Ayonga Hereid

文献摘要

被引文献

相似文献

提出了一种新的强化学习(RL)框架来设计三维双足运动的串级反馈控制策略。现有的RL算法通常是以端到端的方式训练的,或者依赖于一些参考关节或任务空间轨迹的先验知识。与这些研究不同的是,我们提出了一种策略结构,将两足行走问题分解为两个模块,其中包括来自行走动力学本质的物理见解和成熟的用于3D两足步行的混合零动力学方法。因此,整个RL框架具有几个关键优势,包括轻量级网络结构、样本效率和对先验知识的较少依赖。该方法从零开始学习稳定、健壮的步行步态,使控制器能够实现全方位行走,精确跟踪期望速度和航向角度。习得的政策也有力地对抗施加在躯干上的各种敌对力量,并盲目地走在一系列具有挑战性和无结构的地形上。这些结果表明,所提出的串级反馈控制策略适用于室内和室外环境中的三维双足机器人导航。
This paper presents a novel reinforcement learning (RL) framework to design cascade feedback control policies for 3D bipedal locomotion. Existing RL algorithms are often trained in an end-to-end manner or rely on prior knowledge of some reference joint or task space trajectories. Unlike these studies, we propose a policy structure that decouples the bipedal locomotion problem into two modules that incorporate the physical insights from the nature of the walking dynamics and the well-established Hybrid Zero Dynamics approach for 3D bipedal walking. As a result, the overall RL framework has several key advantages, including lightweight network structure, sample efficiency, and less dependence on prior knowledge. The proposed solution learns stable and robust walking gaits from scratch and allows the controller to realize omnidirectional walking with accurate tracking of the desired velocity and heading angle. The learned policies also perform robustly against various adversarial forces applied to the torso and walking blindly on a series of challenging and unstructured terrains. These results demonstrate that the proposed cascade feedback control policy is suitable for navigation of 3D bipedal robots in indoor and outdoor environments.