Demonstration-Guided Reinforcement Learning with Learned Skills

Demonstration-Guided Reinforcement Learning with Learned Skills
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Karl Pertsch;Youngwoon Lee;Yue Wu;Joseph J. Lim
Karl Pertsch;Youngwoon Lee;Yue Wu;Joseph J. Lim
中科院分区:
其他
文献类型:
--
作者:
Karl Pertsch;Youngwoon Lee;Yue Wu;Joseph J. Lim

文献摘要

被引文献

相似文献

演示引导强化学习(RL)是一种通过利用奖励反馈和一组目标任务演示来学习复杂行为的很有前途的方法。以前的演示引导RL方法将每个新任务视为一个独立的学习问题,并试图一步一步地遵循所提供的演示,类似于人类试图通过跟随演示者的确切肌肉动作来模仿完全看不见的行为。当然,这样的学习将是缓慢的,但通常新的行为并不是完全看不见的:它们与我们之前学习的行为共享子任务。在这项工作中,我们的目标是利用这种共享子任务结构来提高演示导引RL的效率。我们首先从从许多任务中收集的先前经验的大型离线数据集中学习一组可重用的技能。然后,我们提出了基于技能的学习与演示(SkiLD),这是一种用于演示制导的RL的算法,它通过遵循演示的技能而不是原始的操作来有效地利用提供的演示,从而比以前的演示制导的RL方法有实质性的性能改进。我们在长视距迷宫导航和复杂的机器人操作任务中验证了该方法的有效性。
Demonstration-guided reinforcement learning (RL) is a promising approach for learning complex behaviors by leveraging both reward feedback and a set of target task demonstrations. Prior approaches for demonstration-guided RL treat every new task as an independent learning problem and attempt to follow the provided demonstrations step-by-step, akin to a human trying to imitate a completely unseen behavior by following the demonstrator's exact muscle movements. Naturally, such learning will be slow, but often new behaviors are not completely unseen: they share subtasks with behaviors we have previously learned. In this work, we aim to exploit this shared subtask structure to increase the efficiency of demonstration-guided RL. We first learn a set of reusable skills from large offline datasets of prior experience collected across many tasks. We then propose Skill-based Learning with Demonstrations (SkiLD), an algorithm for demonstration-guided RL that efficiently leverages the provided demonstrations by following the demonstrated skills instead of the primitive actions, resulting in substantial performance improvements over prior demonstration-guided RL approaches. We validate the effectiveness of our approach on long-horizon maze navigation and complex robot manipulation tasks.