Learning Memory-Based Control for Human-Scale Bipedal Locomotion

Learning Memory-Based Control for Human-Scale Bipedal Locomotion
复制标题

DOI:
10.15607/rss.2020.xvi.031
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
J. Siekmann;S. Valluri;Jeremy Dao;Lorenzo Bermillo;Helei Duan;Alan Fern;J. Hurst
J. Siekmann;S. Valluri;Jeremy Dao;Lorenzo Bermillo;Helei Duan;Alan Fern;J. Hurst
中科院分区:
其他
文献类型:
--
作者:
J. Siekmann;S. Valluri;Jeremy Dao;Lorenzo Bermillo;Helei Duan;Alan Fern;J. Hurst

文献摘要

被引文献

相似文献

控制一个非静态稳定的伺服电机是一个困难的问题,主要是由于复杂的混合动力学。最近的工作已经证明了强化学习(RL)的有效性,基于模拟的神经网络控制器的训练,成功地转移到真实的两足动物。然而,现有的工作主要使用简单的无记忆网络架构,即使更复杂的架构,如包括内存的架构,通常会在其他RL域中产生上级性能。在这项工作中,我们考虑了用于模拟到真实的运动的递归神经网络(RNN),允许学习使用内部记忆来建模重要物理特性的策略。我们表明,虽然RNN能够在模拟中显着优于无记忆策略,但由于对模拟物理的过拟合,它们在真实的模型上没有表现出上级行为,除非使用动态随机化进行训练以防止过拟合;这导致了持续更好的模拟到真实的传输。我们还表明,RNN可以使用它们学习的记忆状态,通过将动态参数编码到记忆中来执行在线系统识别。
Controlling a non-statically stable biped is a difficult problem largely due to the complex hybrid dynamics involved. Recent work has demonstrated the effectiveness of reinforcement learning (RL) for simulation-based training of neural network controllers that successfully transfer to real bipeds. The existing work, however, has primarily used simple memoryless network architectures, even though more sophisticated architectures, such as those including memory, often yield superior performance in other RL domains. In this work, we consider recurrent neural networks (RNNs) for sim-to-real biped locomotion, allowing for policies that learn to use internal memory to model important physical properties. We show that while RNNs are able to significantly outperform memoryless policies in simulation, they do not exhibit superior behavior on the real biped due to overfitting to the simulation physics unless trained using dynamics randomization to prevent overfitting; this leads to consistently better sim-to-real transfer. We also show that RNNs could use their learned memory states to perform online system identification by encoding parameters of the dynamics into memory.