Learning Spring Mass Locomotion: Guiding Policies With a Reduced-Order Model

Learning Spring Mass Locomotion: Guiding Policies With a Reduced-Order Model
复制标题

DOI:
10.1109/lra.2021.3066833
复制
发表时间:
2021-04-01
影响因子:
5.2
通讯作者:
Hurst, Jonathan
Hurst, Jonathan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Green, Kevin;Godse, Yesh;Hurst, Jonathan

文献摘要

被引文献

相似文献

在这封信中,我们描述了一种方法来实现物理机器人的动态腿部运动,它结合了现有的控制方法与强化学习。具体来说,我们的目标是一个控制层次结构,其中最高级别的行为计划通过降阶模型,它描述了腿运动的基本物理,和较低级别的控制器利用学习的政策,可以弥合理想化的,简单的模型和复杂的,全阶机器人之间的差距差距。高级规划器可以使用环境的模型并且是特定于任务的,而低级学习控制器可以执行各种各样的动作,以便它适用于许多不同的任务。在这封信中,我们描述了这种学习动态步行控制器,并表明,从降阶模型的步行运动范围可以用作命令和学习策略的主要训练信号。由此产生的策略不试图天真地跟踪运动(如传统的轨迹跟踪控制器将),而是平衡即时运动跟踪与长期稳定性。由此产生的控制器上演示了人的规模,不受约束,不受约束的双足机器人的速度高达1.2米/秒。这封信建立了一个通用的,动态学习步行控制器,可以应用于许多不同的任务的基础。
In this letter, we describe an approach to achieve dynamic legged locomotion on physical robots which combines existing methods for control with reinforcement learning. Specifically, our goal is a control hierarchy in which highest-level behaviors are planned through reduced-order models, which describe the fundamental physics of legged locomotion, and lower level controllers utilize a learned policy that can bridge the gap between the idealized, simple model and the complex, full order robot. The high-level planner can use a model of the environment and be task specific, while the low-level learned controller can execute a wide range of motions so that it applies to many different tasks. In this letter, we describe this learned dynamic walking controller and show that a range of walking motions from reduced-order models can be used as the command and primary training signal for learned policies. The resulting policies do not attempt to naively track the motion (as a traditional trajectory tracking controller would) but instead balance immediate motion tracking with long term stability. The resulting controller is demonstrated on a human scale, unconstrained, untethered bipedal robot at speeds up to 1.2 m/s. This letter builds the foundation of a generic, dynamic learned walking controller that can be applied to many different tasks.