Masked World Models for Visual Control

Masked World Models for Visual Control
复制标题

DOI:
10.48550/arxiv.2206.14244
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Younggyo Seo;Danijar Hafner;Hao Liu;Fangchen Liu;Stephen James;Kimin Lee;P. Abbeel
Younggyo Seo;Danijar Hafner;Hao Liu;Fangchen Liu;Stephen James;Kimin Lee;P. Abbeel
中科院分区:
其他
文献类型:
--
作者:
Younggyo Seo;Danijar Hafner;Hao Liu;Fangchen Liu;Stephen James;Kimin Lee;P. Abbeel

文献摘要

被引文献

相似文献

基于视觉模型的强化学习(RL)有可能使机器人能够从视觉观察中进行样本高效学习。然而,目前的方法通常是端到端地训练单个模型来学习视觉表示和动态,这使得很难准确地建模机器人和小物体之间的交互。在这项工作中,我们引入了一个基于视觉模型的RL框架,该框架将视觉表示学习和动态学习相结合。具体来说,我们用卷积层和视觉变换器(ViT)训练自动编码器,以重建给定掩蔽卷积特征的像素,并学习一个对自动编码器的表示进行操作的潜在动态模型。此外,为了对任务相关信息进行编码,我们为自动编码器引入了辅助奖励预测目标。我们使用从环境交互中收集的在线样本不断更新自动编码器和动力学模型。我们证明了我们的解耦方法在Meta-world和RLBench的各种视觉机器人任务上实现了最先进的性能,例如,我们在Meta-world的50个视觉机器人操作任务上取得了81.7%的成功率,而基线为67.9%。代码可在项目网站上获得:https://sites.google.com/view/mwm-rl。
Visual model-based reinforcement learning (RL) has the potential to enable sample-efficient robot learning from visual observations. Yet the current approaches typically train a single model end-to-end for learning both visual representations and dynamics, making it difficult to accurately model the interaction between robots and small objects. In this work, we introduce a visual model-based RL framework that decouples visual representation learning and dynamics learning. Specifically, we train an autoencoder with convolutional layers and vision transformers (ViT) to reconstruct pixels given masked convolutional features, and learn a latent dynamics model that operates on the representations from the autoencoder. Moreover, to encode task-relevant information, we introduce an auxiliary reward prediction objective for the autoencoder. We continually update both autoencoder and dynamics model using online samples collected from environment interaction. We demonstrate that our decoupling approach achieves state-of-the-art performance on a variety of visual robotic tasks from Meta-world and RLBench, e.g., we achieve 81.7% success rate on 50 visual robotic manipulation tasks from Meta-world, while the baseline achieves 67.9%. Code is available on the project website: https://sites.google.com/view/mwm-rl.