Reset-free Trial-and-Error Learning for Robot Damage Recovery

Reset-free Trial-and-Error Learning for Robot Damage Recovery
复制标题

DOI:
10.1016/j.robot.2017.11.010
复制
发表时间:
2018-02-01
影响因子:
4.3
通讯作者:
Mouret, Jean-Baptiste
Mouret, Jean-Baptiste
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chatzilygeroudis, Konstantinos;Vassiliades, Vassilis;Mouret, Jean-Baptiste

文献摘要

被引文献

相似文献

硬件故障的高概率阻碍了许多先进的机器人(例如,腿式机器人)在真实世界的情况下(例如,灾后救援)被自信地部署。机器人可以通过反复试验来适应,而不是试图诊断故障,以便能够完成它们的任务。在这种情况下,损伤恢复可以看作是一个强化学习(RL)问题。然而,用于机器人的最好的RL算法要求机器人和环境在每一集之后被重置到初始状态,即机器人不是自主学习的。此外,大多数用于机器人学的RL方法不能很好地扩展到复杂的机器人(例如,行走机器人),并且要么根本不能使用,要么需要太长的时间来收敛到解决方案(例如,学习时间)。本文提出了一种新的学习算法“无重置试错法”(RTE),该算法(1)通过使用完整机器人的动力学模拟器预先生成数百种可能的行为来打破复杂性,(2)允许复杂的机器人在完成任务并考虑环境的同时快速恢复受损。我们在一个模拟的轮式机器人、一个模拟的六足机器人和一个真实的六足步行机器人上对我们的算法进行了评估,这些机器人在几个方面都受到了损害(例如,一条腿缺失、一条腿缩短、电机故障等)。其目标是达到竞技场上的一系列目标。我们的实验表明,机器人可以在没有任何人为干预的情况下,在有障碍物的环境中恢复大部分运动能力。(C)2017年作者。Elsevier B.V.出版。这是一篇基于CC by License(http://creativecommons.orgilicenses/by/4.0/).的开放获取文章
The high probability of hardware failures prevents many advanced robots (e.g., legged robots) from being confidently deployed in real-world situations (e.g., post-disaster rescue). Instead of attempting to diagnose the failures, robots could adapt by trial-and-error in order to be able to complete their tasks. In this situation, damage recovery can be seen as a Reinforcement Learning (RL) problem. However, the best RL algorithms for robotics require the robot and the environment to be reset to an initial state after each episode, that is, the robot is not learning autonomously. In addition, most of the RL methods for robotics do not scale well with complex robots (e.g., walking robots) and either cannot be used at all or take too long to converge to a solution (e.g., hours of learning). In this paper, we introduce a novel learning algorithm called "Reset-free Trial-and-Error" (RTE) that (1) breaks the complexity by pre-generating hundreds of possible behaviors with a dynamics simulator of the intact robot, and (2) allows complex robots to quickly recover from damage while completing their tasks and taking the environment into account. We evaluate our algorithm on a simulated wheeled robot, a simulated six-legged robot, and a real six-legged walking robot that are damaged in several ways (e.g., a missing leg, a shortened leg, faulty motor, etc.) and whose objective is to reach a sequence of targets in an arena. Our experiments show that the robots can recover most of their locomotion abilities in an environment with obstacles, and without any human intervention. (C) 2017 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY license (http://creativecommons.orgilicenses/by/4.0/).