DayDreamer: World Models for Physical Robot Learning

DayDreamer: World Models for Physical Robot Learning
复制标题

DOI:
10.48550/arxiv.2206.14176
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Philipp Wu;Alejandro Escontrela;Danijar Hafner;Ken Goldberg;P. Abbeel
Philipp Wu;Alejandro Escontrela;Danijar Hafner;Ken Goldberg;P. Abbeel
中科院分区:
其他
文献类型:
--
作者:
Philipp Wu;Alejandro Escontrela;Danijar Hafner;Ken Goldberg;P. Abbeel

文献摘要

被引文献

相似文献

为了在复杂环境中完成任务,机器人需要从经验中学习。深度强化学习是机器人学习的一种常用方法,但需要大量的试错来学习,这限制了它在现实世界中的应用。因此,机器人学习的许多进展都依赖于模拟器。另一方面,在模拟器中学习无法捕捉现实世界的复杂性,容易出现模拟器不准确的情况,而且所产生的行为无法适应现实世界的变化。Dreamer算法最近在通过在学习到的世界模型中进行规划,从少量交互中学习方面显示出巨大的潜力,在视频游戏中优于单纯的强化学习。学习一个世界模型来预测潜在行动的结果能够在想象中进行规划,减少在现实环境中所需的试错量。然而,尚不清楚Dreamer是否能促进物理机器人更快地学习。在本文中,我们将Dreamer应用于4种机器人,在没有模拟器的情况下直接在现实世界中进行在线学习。Dreamer训练一个四足机器人从仰卧状态翻滚、站立并行走,而且从头开始且无需重置,仅用1小时。然后我们推动机器人,发现Dreamer能在10分钟内适应,以承受干扰,或者快速翻滚并重新站立起来。在两个不同的机械臂上,Dreamer学会直接从相机图像和稀疏奖励中抓取和放置多个物体,接近人类的表现。在一个轮式机器人上,Dreamer学会纯粹从相机图像导航到目标位置,自动解决机器人方向的模糊性。在所有实验中使用相同的超参数,我们发现Dreamer能够在现实世界中进行在线学习,建立了一个强大的基准。我们发布我们的基础设施,以便未来将世界模型应用于机器人学习。
To solve tasks in complex environments, robots need to learn from experience. Deep reinforcement learning is a common approach to robot learning but requires a large amount of trial and error to learn, limiting its deployment in the physical world. As a consequence, many advances in robot learning rely on simulators. On the other hand, learning inside of simulators fails to capture the complexity of the real world, is prone to simulator inaccuracies, and the resulting behaviors do not adapt to changes in the world. The Dreamer algorithm has recently shown great promise for learning from small amounts of interaction by planning within a learned world model, outperforming pure reinforcement learning in video games. Learning a world model to predict the outcomes of potential actions enables planning in imagination, reducing the amount of trial and error needed in the real environment. However, it is unknown whether Dreamer can facilitate faster learning on physical robots. In this paper, we apply Dreamer to 4 robots to learn online and directly in the real world, without simulators. Dreamer trains a quadruped robot to roll off its back, stand up, and walk from scratch and without resets in only 1 hour. We then push the robot and find that Dreamer adapts within 10 minutes to withstand perturbations or quickly roll over and stand back up. On two different robotic arms, Dreamer learns to pick and place multiple objects directly from camera images and sparse rewards, approaching human performance. On a wheeled robot, Dreamer learns to navigate to a goal position purely from camera images, automatically resolving ambiguity about the robot orientation. Using the same hyperparameters across all experiments, we find that Dreamer is capable of online learning in the real world, establishing a strong baseline. We release our infrastructure for future applications of world models to robot learning.