Learning Navigation Subroutines from Egocentric Videos

Learning Navigation Subroutines from Egocentric Videos
复制标题

从以自我为中心的视频中学习导航子程序

DOI:
--
复制
发表时间:
2019
期刊:
Conference on Robot Learning
影响因子:
--
通讯作者:
J. Malik
J. Malik
中科院分区:
--
文献类型:
--
作者:
Ashish Kumar;Saurabh Gupta;J. Malik

文献摘要

被引文献

相似文献

在更高的抽象级别而不是低级别的扭矩上进行规划可以提高强化学习中的样本效率以及经典规划中的计算效率。我们提出了一种方法,可以从执行任务的专家以自我为中心的视频数据中学习这种分层抽象或子例程。我们在少量随机交互数据上学习自监督逆模型,以用代理动作伪标记专家以自我为中心的视频。通过学习潜在的意图条件策略,从这些伪标记视频中获取视觉运动子程序,该策略根据相应的图像观察来预测推断的伪动作。我们在导航背景下展示了我们提出的方法,并表明我们可以成功地从被动的自我中心视频中学习一致且多样化的视觉运动子程序。我们通过将所获得的视觉运动子程序按原样用于探索,并作为分层强化学习框架中的子策略来实现点目标和语义目标,从而展示了它们的实用性。我们还通过将子程序部署在真实的机器人平台上来演示它们在现实世界中的行为。项目网站:这个 https URL。
Planning at a higher level of abstraction instead of low level torques improves the sample efficiency in reinforcement learning, and computational efficiency in classical planning. We propose a method to learn such hierarchical abstractions, or subroutines from egocentric video data of experts performing tasks. We learn a self-supervised inverse model on small amounts of random interaction data to pseudo-label the expert egocentric videos with agent actions. Visuomotor subroutines are acquired from these pseudo-labeled videos by learning a latent intent-conditioned policy that predicts the inferred pseudo-actions from the corresponding image observations. We demonstrate our proposed approach in context of navigation, and show that we can successfully learn consistent and diverse visuomotor subroutines from passive egocentric videos. We demonstrate the utility of our acquired visuomotor subroutines by using them as is for exploration, and as sub-policies in a hierarchical RL framework for reaching point goals and semantic goals. We also demonstrate behavior of our subroutines in the real world, by deploying them on a real robotic platform. Project website: this https URL.