A Survey on Policy Search Algorithms for Learning Robot Controllers in a Handful of Trials

A Survey on Policy Search Algorithms for Learning Robot Controllers in a Handful of Trials
复制标题

DOI:
10.1109/tro.2019.2958211
复制
发表时间:
2020-04-01
影响因子:
7.8
通讯作者:
Mouret, Jean-Baptiste
Mouret, Jean-Baptiste
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chatzilygeroudis, Konstantinos;Vassiliades, Vassilis;Mouret, Jean-Baptiste

文献摘要

被引文献

相似文献

大多数策略搜索(PS)算法需要数千个训练集才能找到有效的策略,这对于物理机器人来说通常是不可行的。这篇调查文章关注的是极端的另一端:机器人如何在几次试验(十几次)和几分钟内适应?通过类比“大数据”这个词,我们将这一挑战称为“微数据强化学习”。“在这篇文章中,我们展示了第一个策略是利用关于政策结构的先验知识(例如,动态移动原语),策略参数(例如,演示),或动态(例如,模拟器)。第二种策略是创建预期回报的数据驱动的代理模型(例如,贝叶斯优化)或动态模型(例如,基于模型的PS),使得策略优化器查询模型而不是真实的系统。总的来说,所有成功的微数据算法通过改变模型和先验知识的种类将这两种策略联合收割机结合起来。目前的科学挑战主要围绕着扩展到复杂的机器人,设计通用的先验知识,以及优化计算时间。
Most policy search (PS) algorithms require thousands of training episodes to find an effective policy, which is often infeasible with a physical robot. This survey article focuses on the extreme other end of the spectrum: how can a robot adapt with only a handful of trials (a dozen) and a few minutes? By analogy with the word "big-data," we refer to this challenge as "micro-data reinforcement learning." In this article, we show that a first strategy is to leverage prior knowledge on the policy structure (e.g., dynamic movement primitives), on the policy parameters (e.g., demonstrations), or on the dynamics (e.g., simulators). A second strategy is to create data-driven surrogate models of the expected reward (e.g., Bayesian optimization) or the dynamical model (e.g., model-based PS), so that the policy optimizer queries the model instead of the real system. Overall, all successful micro-data algorithms combine these two strategies by varying the kind of model and prior knowledge. The current scientific challenges essentially revolve around scaling up to complex robots, designing generic priors, and optimizing the computing time.