Masked Autoencoding for Scalable and Generalizable Decision Making

Masked Autoencoding for Scalable and Generalizable Decision Making
复制标题

DOI:
10.48550/arxiv.2211.12740
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Fangchen Liu;Hao Liu;Aditya Grover;P. Abbeel
Fangchen Liu;Hao Liu;Aditya Grover;P. Abbeel
中科院分区:
其他
文献类型:
--
作者:
Fangchen Liu;Hao Liu;Aditya Grover;P. Abbeel

文献摘要

相似文献

我们有兴趣学习用于强化学习的可扩展代理,这些代理可以从类似于当前大型视觉和语言模型的大规模、多样化序列数据中学习。为此,本文提出了掩蔽决策预测(MaskDP),这是一种用于强化学习(RL)和行为克隆(BC)的简单且可扩展的自监督预训练方法。在我们的 MaskDP 方法中,我们采用掩码自动编码器(MAE)来描述状态动作轨迹,其中我们随机掩码状态和动作标记并重建丢失的数据。通过这样做,模型需要推断隐藏的状态和动作并提取有关动态的信息。我们发现,屏蔽输入序列的不同比例显着有助于学习更好的模型,该模型可以很好地推广到多个下游任务。在我们的实证研究中,我们发现 MaskDP 模型获得了零样本迁移到新的 BC 任务的能力,例如单个和多个目标达成,并且它可以从一些示例转换中零样本推断技能。此外,MaskDP 可以很好地迁移到离线 RL,并显示出有希望的扩展行为。到模型尺寸。它适合于数据高效的微调,与基于自回归预训练的现有方法相比,可以获得有竞争力的结果。
We are interested in learning scalable agents for reinforcement learning that can learn from large-scale, diverse sequential data similar to current large vision and language models. To this end, this paper presents masked decision prediction (MaskDP), a simple and scalable self-supervised pretraining method for reinforcement learning (RL) and behavioral cloning (BC). In our MaskDP approach, we employ a masked autoencoder (MAE) to state-action trajectories, wherein we randomly mask state and action tokens and reconstruct the missing data. By doing so, the model is required to infer masked-out states and actions and extract information about dynamics. We find that masking different proportions of the input sequence significantly helps with learning a better model that generalizes well to multiple downstream tasks. In our empirical study, we find that a MaskDP model gains the capability of zero-shot transfer to new BC tasks, such as single and multiple goal reaching, and it can zero-shot infer skills from a few example transitions. In addition, MaskDP transfers well to offline RL and shows promising scaling behavior w.r.t. to model size. It is amenable to data-efficient finetuning, achieving competitive results with prior methods based on autoregressive pretraining.