Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills

Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills
复制标题

DOI:
--
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
Yevgen Chebotar;Karol Hausman;Yao Lu-;Ted Xiao;Dmitry Kalashnikov;Jacob Varley;A. Irpan;Benjamin Eysenbach;Ryan C. Julian;Chelsea Finn;S. Levine
Yevgen Chebotar;Karol Hausman;Yao Lu-;Ted Xiao;Dmitry Kalashnikov;Jacob Varley;A. Irpan;Benjamin Eysenbach;Ryan C. Julian;Chelsea Finn;S. Levine
中科院分区:
其他
文献类型:
--
作者:
Yevgen Chebotar;Karol Hausman;Yao Lu-;Ted Xiao;Dmitry Kalashnikov;Jacob Varley;A. Irpan;Benjamin Eysenbach;Ryan C. Julian;Chelsea Finn;S. Levine

文献摘要

被引文献

相似文献

我们考虑从之前收集的离线数据中学习有用的机器人技能的问题,而无需访问手动指定的奖励或额外的在线探索,这种设置对于通过重用过去的机器人数据来扩展机器人学习变得越来越重要。特别是,我们提出了通过学习达到给定数据集中的任何目标状态来学习对环境的功能理解的目标。我们采用目标条件Q学习与后见之明重新标记,并开发了几种技术,使培训在一个特别具有挑战性的离线设置。我们发现,我们的方法可以在高维相机图像上操作,并在真实的机器人上学习各种技能,这些机器人可以推广到以前看不见的场景和物体。我们还表明,我们的方法可以通过目标链学习在多个事件中实现长期目标,并通过预训练或辅助目标学习丰富的表示,可以帮助下游任务。我们的实验视频可以在https://actionable-models.github.io上找到
We consider the problem of learning useful robotic skills from previously collected offline data without access to manually specified rewards or additional online exploration, a setting that is becoming increasingly important for scaling robot learning by reusing past robotic data. In particular, we propose the objective of learning a functional understanding of the environment by learning to reach any goal state in a given dataset. We employ goal-conditioned Q-learning with hindsight relabeling and develop several techniques that enable training in a particularly challenging offline setting. We find that our method can operate on high-dimensional camera images and learn a variety of skills on real robots that generalize to previously unseen scenes and objects. We also show that our method can learn to reach long-horizon goals across multiple episodes through goal chaining, and learn rich representations that can help with downstream tasks through pre-training or auxiliary objectives. The videos of our experiments can be found at https://actionable-models.github.io