Adaptive Prior Selection for Repertoire-Based Online Adaptation in Robotics.

Adaptive Prior Selection for Repertoire-Based Online Adaptation in Robotics.
复制标题

DOI:
10.3389/frobt.2019.00151
复制
发表时间:
2019
影响因子:
3.4
通讯作者:
Mouret JB
Mouret JB
中科院分区:
其他
文献类型:
--
作者:
Kaushik R;Desreumaux P;Mouret JB

文献摘要

参考文献

被引文献

相似文献

基于指令集的学习是一种基于两步过程的数据高效适应方法,其中(1)在模拟中学习大量且多样化的策略,(2)规划或学习算法根据当前情况(例如,损坏的机器人、新物体等)选择最合适的策略。在本文中,我们放宽了先前作品的假设,即单个剧目足以进行改编。相反,我们为许多不同的情况生成曲目(例如,缺少一条腿、在不同的楼层等),并让我们的算法选择最有用的先验。我们的主要贡献是一种算法,APROL(基于曲目的在线学习的自适应先验选择),当机器人没有关于当前情况的信息时,通过结合这些先验来规划下一步行动。我们在两个模拟任务上评估 APROL:(1)用机械臂推动各种形状和大小的未知物体,以及(2)用损坏的六足机器人完成目标达成任务。我们与“免重置试错”(RTE)和各种基于单一指令集的基线进行比较。结果表明,APROL 以比基线更少的交互时间解决了这两项任务。此外,我们还在一个真实的、受损的六足动物上演示了 APROL,它可以快速学会选择补偿策略,通过避开路径中的障碍来实现目标。
Repertoire-based learning is a data-efficient adaptation approach based on a two-step process in which (1) a large and diverse set of policies is learned in simulation, and (2) a planning or learning algorithm chooses the most appropriate policies according to the current situation (e.g., a damaged robot, a new object, etc.). In this paper, we relax the assumption of previous works that a single repertoire is enough for adaptation. Instead, we generate repertoires for many different situations (e.g., with a missing leg, on different floors, etc.) and let our algorithm selects the most useful prior. Our main contribution is an algorithm, APROL (Adaptive Prior selection for Repertoire-based Online Learning) to plan the next action by incorporating these priors when the robot has no information about the current situation. We evaluate APROL on two simulated tasks: (1) pushing unknown objects of various shapes and sizes with a robotic arm and (2) a goal reaching task with a damaged hexapod robot. We compare with “Reset-free Trial and Error” (RTE) and various single repertoire-based baselines. The results show that APROL solves both the tasks in less interaction time than the baselines. Additionally, we demonstrate APROL on a real, damaged hexapod that quickly learns to pick compensatory policies to reach a goal by avoiding obstacles in the path.
DOI: 10.1109/tpami.2013.218
发表时间: 2015-02-01
影响因子: 23.6
作者:
Deisenroth, Marc Peter;Fox, Dieter;Rasmussen, Carl Edward
通讯作者: Rasmussen, Carl Edward
DOI: 10.1023/a:1013689704352
发表时间: 2002-01-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Auer, P;Cesa-Bianchi, N;Fischer, P
通讯作者: Fischer, P
DOI: 10.1016/j.robot.2017.11.010
发表时间: 2018-02-01
影响因子: 4.3
作者:
Chatzilygeroudis, Konstantinos;Vassiliades, Vassilis;Mouret, Jean-Baptiste
通讯作者: Mouret, Jean-Baptiste
DOI: 10.1109/tevc.2017.2704781
发表时间: 2018-04-01
影响因子: 14.3
作者:
Cully, Antoine;Demiris, Yiannis
通讯作者: Demiris, Yiannis
DOI: 10.1162/evco_a_00143
发表时间: 2016-03-01
影响因子: 6.8
作者:
Cully, A.;Mouret, J. -B.
通讯作者: Mouret, J. -B.