Learning to Generalize Across Long-Horizon Tasks from Human Demonstrations

Learning to Generalize Across Long-Horizon Tasks from Human Demonstrations
复制标题

从人类演示中学习泛化长期任务

DOI:
--
复制
发表时间:
2020
期刊:
Robotics: Science and Systems
影响因子:
--
通讯作者:
Li Fei
Li Fei
中科院分区:
--
文献类型:
--
作者:
Ajay Mandlekar;Danfei Xu;Roberto Martín;S. Savarese;Li Fei

文献摘要

被引文献

相似文献

模仿学习是在真实的世界中训练机器人策略的有效且安全的技术,因为它不依赖于昂贵的随机探索过程。然而,由于缺乏探索,学习政策,推广以外的表现出的行为仍然是一个开放的挑战。我们提出了一种新的模仿学习框架,使机器人1)学习复杂的真实的世界的操作任务,有效地从少量的人类示范,2)合成新的行为不包含在收集的示范。我们的关键见解是,多任务域往往呈现出一种潜在的结构,其中不同任务的演示轨迹在状态空间的公共区域相交。我们提出了泛化通过模仿(GTI),一个两阶段的离线模仿学习算法,利用这种交叉结构来训练目标导向的政策,推广到看不见的开始和目标状态组合。在GTI的第一阶段,我们训练了一个随机策略,该策略利用轨迹交叉点来组合来自不同演示轨迹的行为。在GTI的第二阶段,我们收集了一个小的推出从第一阶段的无条件随机政策,并训练目标导向代理推广到新的开始和目标配置。我们验证GTI在两个模拟域和一个具有挑战性的长期视野机器人操作域在真实的世界。其他结果和视频可在此https URL。
Imitation learning is an effective and safe technique to train robot policies in the real world because it does not depend on an expensive random exploration process. However, due to the lack of exploration, learning policies that generalize beyond the demonstrated behaviors is still an open challenge. We present a novel imitation learning framework to enable robots to 1) learn complex real world manipulation tasks efficiently from a small number of human demonstrations, and 2) synthesize new behaviors not contained in the collected demonstrations. Our key insight is that multi-task domains often present a latent structure, where demonstrated trajectories for different tasks intersect at common regions of the state space. We present Generalization Through Imitation (GTI), a two-stage offline imitation learning algorithm that exploits this intersecting structure to train goal-directed policies that generalize to unseen start and goal state combinations. In the first stage of GTI, we train a stochastic policy that leverages trajectory intersections to have the capacity to compose behaviors from different demonstration trajectories together. In the second stage of GTI, we collect a small set of rollouts from the unconditioned stochastic policy of the first stage, and train a goal-directed agent to generalize to novel start and goal configurations. We validate GTI in both simulated domains and a challenging long-horizon robotic manipulation domain in the real world. Additional results and videos are available at this https URL .