Unsupervised Perceptual Rewards for Imitation Learning

Unsupervised Perceptual Rewards for Imitation Learning
复制标题

DOI:
10.15607/rss.2017.xiii.050
复制
发表时间:
2016-12
期刊:
ArXiv
影响因子:
--
通讯作者:
P. Sermanet;Kelvin Xu;S. Levine
P. Sermanet;Kelvin Xu;S. Levine
中科院分区:
其他
文献类型:
--
作者:
P. Sermanet;Kelvin Xu;S. Levine

文献摘要

被引文献

相似文献

奖励函数设计和探索时间可以说是在现实世界中部署强化学习(RL)代理的最大障碍。在许多现实世界的任务中,设计奖励函数需要大量的手工工程,并且通常需要安装额外的传感器来测量任务是否已成功执行。此外,许多有趣的任务包含多个必须按顺序执行的隐式中间步骤。即使最终结果可以衡量,也不一定会提供有关这些中间步骤的反馈。为了解决这些问题,我们建议利用深度模型学习的中间视觉表示的抽象能力,从少量演示中快速推断出感知奖励函数。我们提出了一种方法,能够仅从少数演示序列中识别任务的关键中间步骤,并自动识别用于识别这些步骤的最具辨别力的特征。该方法利用预训练深度模型中的特征,但不需要任何明确的子目标规范。然后,强化学习代理可以使用生成的奖励函数来学习如何在现实环境中执行任务。为了评估学习到的奖励,我们提出了两个现实世界任务的定性结果以及针对人类设计的奖励函数的定量评估。我们还表明,我们的方法可以用于使用真实的机器人来学习真实世界的开门技能,即使用于奖励学习的演示是由人类用自己的手提供的。据我们所知,这些是第一个结果,表明可以直接学习复杂的机器人操作技能,而无需从人类执行任务的视频中进行监督标签。补充材料和数据可在此 https URL 获取
Reward function design and exploration time are arguably the biggest obstacles to the deployment of reinforcement learning (RL) agents in the real world. In many real-world tasks, designing a reward function takes considerable hand engineering and often requires additional sensors to be installed just to measure whether the task has been executed successfully. Furthermore, many interesting tasks consist of multiple implicit intermediate steps that must be executed in sequence. Even when the final outcome can be measured, it does not necessarily provide feedback on these intermediate steps. To address these issues, we propose leveraging the abstraction power of intermediate visual representations learned by deep models to quickly infer perceptual reward functions from small numbers of demonstrations. We present a method that is able to identify key intermediate steps of a task from only a handful of demonstration sequences, and automatically identify the most discriminative features for identifying these steps. This method makes use of the features in a pre-trained deep model, but does not require any explicit specification of sub-goals. The resulting reward functions can then be used by an RL agent to learn to perform the task in real-world settings. To evaluate the learned reward, we present qualitative results on two real-world tasks and a quantitative evaluation against a human-designed reward function. We also show that our method can be used to learn a real-world door opening skill using a real robot, even when the demonstration used for reward learning is provided by a human using their own hand. To our knowledge, these are the first results showing that complex robotic manipulation skills can be learned directly and without supervised labels from a video of a human performing the task. Supplementary material and data are available at this https URL