Data-Driven Latent Space Representation for Robust Bipedal Locomotion Learning

Data-Driven Latent Space Representation for Robust Bipedal Locomotion Learning
复制标题

DOI:
10.48550/arxiv.2309.15740
复制
发表时间:
2023-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Guillermo A. Castillo;Bowen Weng;Wei Zhang;Ayonga Hereid
Guillermo A. Castillo;Bowen Weng;Wei Zhang;Ayonga Hereid
中科院分区:
其他
文献类型:
--
作者:
Guillermo A. Castillo;Bowen Weng;Wei Zhang;Ayonga Hereid

文献摘要

相似文献

将数据驱动的状态表示与基于强化学习(RL)的运动策略相结合,提出了一种学习健壮两足步行的新框架。该框架利用一个自动编码器来学习一个低维潜在空间,该空间从现有的运动数据中捕捉到两足动物运动的复杂动力学。然后,这种降维的状态表示被用作训练基于RL的稳健步态策略的状态,从而消除了对启发式状态选择或用于步态规划的模板模型的需要。结果表明,学习的潜变量是解缠的,直接对应于不同的步态或速度,如前进、后退或原地行走。与传统的基于模板模型的方法相比,该框架在仿真方面表现出更好的性能和健壮性。经过训练的策略有效地跟踪了大范围的行走速度,并展示了对未知场景的良好泛化能力。
This paper presents a novel framework for learning robust bipedal walking by combining a data-driven state representation with a Reinforcement Learning (RL) based locomotion policy. The framework utilizes an autoencoder to learn a low-dimensional latent space that captures the complex dynamics of bipedal locomotion from existing locomotion data. This reduced dimensional state representation is then used as states for training a robust RL-based gait policy, eliminating the need for heuristic state selections or the use of template models for gait planning. The results demonstrate that the learned latent variables are disentangled and directly correspond to different gaits or speeds, such as moving forward, backward, or walking in place. Compared to traditional template model-based approaches, our framework exhibits superior performance and robustness in simulation. The trained policy effectively tracks a wide range of walking speeds and demonstrates good generalization capabilities to unseen scenarios.