Online Learning of Unknown Dynamics for Model-Based Controllers in Legged Locomotion

Online Learning of Unknown Dynamics for Model-Based Controllers in Legged Locomotion
复制标题

DOI:
10.1109/lra.2021.3108510
复制
发表时间:
2021-10
影响因子:
5.2
通讯作者:
Yu Sun;Wyatt Ubellacker;Wen-Loong Ma;Xiang Zhang;Changhao Wang;Noel Csomay-Shanklin;M. Tomizuka;K. Sreenath;A. Ames
Yu Sun;Wyatt Ubellacker;Wen-Loong Ma;Xiang Zhang;Changhao Wang;Noel Csomay-Shanklin;M. Tomizuka;K. Sreenath;A. Ames
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yu Sun;Wyatt Ubellacker;Wen-Loong Ma;Xiang Zhang;Changhao Wang;Noel Csomay-Shanklin;M. Tomizuka;K. Sreenath;A. Ames

文献摘要

被引文献

相似文献

当基于模型的控制器的模型不准确地表示真实的世界动态时,其性能会受到严重影响。我们建议学习一个随时间变化的,局部线性残差模型沿着机器人的当前轨迹,以补偿控制器的模型的预测误差。监督学习是在线执行的,因为机器人在未知的环境中运行,使用从其最近的过去收集的数据。我们从理论上研究我们的方法在其一般配方,然后将其应用到一个双足控制器来自全阶动态的虚拟约束,和一个四足控制器来自一个简化模型的接触力。对于模拟,我们的方法始终优于基线和最近的基于学习的方法。我们还在模拟和真实的世界中用12公斤的四足动物进行了实验,其中基线无法使用10公斤的有效载荷行走,但我们的方法成功了。
The performance of a model-based controller can severely suffer when its model inaccurately represents the real world dynamics. We propose to learn a time-varying, locally linear residual model along the robot’s current trajectory, to compensate for the prediction errors of the controller’s model. Supervised learning is performed online, as the robot is running in the unknown environment, using data collected from its immediate past. We theoretically investigate our method in its general formulation, then apply it to a bipedal controller derived from the full-order dynamics of virtual constraints, and a quadrupedal controller derived from a simplified model of contact forces. For a biped in simulation, our method consistently outperforms the baseline and a recent learning-based method. We also experiment with a 12 kg quadruped in simulation and real world, where the baseline fails to walk with 10 kg of payload but our method succeeds.