Gaussian Processes for Data-Efficient Learning in Robotics and Control

Gaussian Processes for Data-Efficient Learning in Robotics and Control
复制标题

DOI:
10.1109/tpami.2013.218
复制
发表时间:
2015-02-01
影响因子:
23.6
通讯作者:
Rasmussen, Carl Edward
Rasmussen, Carl Edward
中科院分区:
计算机科学1区
文献类型:
--
作者:
Deisenroth, Marc Peter;Fox, Dieter;Rasmussen, Carl Edward

文献摘要

被引文献

相似文献

十多年来,自主学习一直是控制和机器人技术的一个有前途的方向,因为数据驱动的学习可以减少工程知识的数量。然而,自主强化学习(RL)方法通常需要与系统的许多交互来学习控制器,这在诸如机器人的真实的系统中是实际限制,其中许多交互可能是不切实际的并且耗时的。为了解决这个问题,目前的学习方法通常需要以专家演示、逼真的模拟器、预先成形的策略或有关底层动态的特定知识的形式的特定于任务的知识。在本文中,我们采用了不同的方法,通过从数据中提取更多信息来加快学习速度。特别是,我们学习的概率,非参数高斯过程过渡模型的系统。通过明确地将模型的不确定性纳入长期规划和控制器学习,我们的方法减少了模型误差的影响,这是基于模型的学习中的一个关键问题。与最先进的RL相比,我们基于模型的策略搜索方法实现了前所未有的学习速度。我们证明了它的适用性自主学习在真实的机器人和控制任务。
Autonomous learning has been a promising direction in control and robotics for more than a decade since data-driven learning allows to reduce the amount of engineering knowledge, which is otherwise required. However, autonomous reinforcement learning (RL) approaches typically require many interactions with the system to learn controllers, which is a practical limitation in real systems, such as robots, where many interactions can be impractical and time consuming. To address this problem, current learning approaches typically require task-specific knowledge in form of expert demonstrations, realistic simulators, pre-shaped policies, or specific knowledge about the underlying dynamics. In this paper, we follow a different approach and speed up learning by extracting more information from data. In particular, we learn a probabilistic, non-parametric Gaussian process transition model of the system. By explicitly incorporating model uncertainty into long-term planning and controller learning our approach reduces the effects of model errors, a key problem in model-based learning. Compared to state-of-the art RL our model-based policy search method achieves an unprecedented speed of learning. We demonstrate its applicability to autonomous learning in real robot and control tasks.