Keep It Simple: Data-Efficient Learning for Controlling Complex Systems With Simple Models

Keep It Simple: Data-Efficient Learning for Controlling Complex Systems With Simple Models
复制标题

DOI:
10.1109/lra.2021.3056368
复制
发表时间:
2021-02
影响因子:
5.2
通讯作者:
Thomas Power;D. Berenson
Thomas Power;D. Berenson
中科院分区:
计算机科学2区
文献类型:
--
作者:
Thomas Power;D. Berenson

文献摘要

被引文献

相似文献

当操纵具有复杂动力学的新对象时,状态表示并不总是可用的,例如,可变形对象。从观察中学习表示和动态都需要大量的数据。我们提出了学习的视觉相似性预测控制(LVSPC),这是一种新的数据有效学习方法,可以从图像中控制具有复杂动态和高维状态空间的系统。LVSPC利用给定的简单模型近似,从该模型可以生成图像观测。我们使用这些图像来训练感知模型,该模型通过在线观察复杂系统来估计简单模型状态。然后,我们使用来自复杂系统的数据来拟合简单模型的参数,并在线了解该模型的不准确之处。最后,我们使用模型预测控制,并将控制器偏离简单模型不准确的区域,从而使控制器不太可靠。我们评估LVSPC两个任务;操纵拴系质量和绳子。我们发现,我们的方法在数据数量级较少的情况下执行最先进的强化学习方法。LVSPC还在真实的机器人上完成了绳索操作任务,仅经过10次试验就有80%的成功率,尽管使用的感知系统只在模拟图像上训练。
When manipulating a novel object with complex dynamics, a state representation is not always available, for example, deformable objects. Learning both a representation and dynamics from observations requires large amounts of data. We propose Learned Visual Similarity Predictive Control (LVSPC), a novel method for data-efficient learning to control systems with complex dynamics and high-dimensional state spaces from images. LVSPC leverages a given simple model approximation from which image observations can be generated. We use these images to train a perception model that estimates the simple model state from observations of the complex system online. We then use data from the complex system to fit the parameters of the simple model and learn where this model is inaccurate, also online. Finally, we use Model Predictive Control and bias the controller away from regions where the simple model is inaccurate and thus where the controller is less reliable. We evaluate LVSPC on two tasks; manipulating a tethered mass and a rope. We find that our method performs comparably to state-of-the-art reinforcement learning methods with an order of magnitude less data. LVSPC also completes the rope manipulation task on a real robot with 80% success rate after only 10 trials, despite using a perception system trained only on images from simulation.