Decision Transformer: Reinforcement Learning via Sequence Modeling

Decision Transformer: Reinforcement Learning via Sequence Modeling
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Lili Chen;Kevin Lu;A. Rajeswaran;Kimin Lee;Aditya Grover;M. Laskin;P. Abbeel;A. Srinivas;
Lili Chen;Kevin Lu;A. Rajeswaran;Kimin Lee;Aditya Grover;M. Laskin;P. Abbeel;A. Srinivas;
中科院分区:
其他
文献类型:
--
作者:
Lili Chen;Kevin Lu;A. Rajeswaran;Kimin Lee;Aditya Grover;M. Laskin;P. Abbeel;A. Srinivas;

文献摘要

被引文献

相似文献

我们介绍了一个将增强学习(RL)作为序列建模问题的框架。这使我们能够借鉴变压器体系结构的简单性和可扩展性,以及语言建模(例如GPT-X和BERT)的相关进展。特别是,我们提出了决策变压器,这是一种将RL作为条件序列建模的架构。与拟合价值功能或计算策略梯度的RL的先前方法不同,决策变压器只需通过利用因果掩盖的变压器来输出最佳动作。通过根据所需的回报(奖励),过去的状态和行动来调节自回旋模型,我们的决策变压器模型可以生成未来的动作,以实现所需的回报。尽管它很简单,但决策变压器匹配或超过了Atari,OpenAI体育馆和钥匙到门任务上最先进的离线RL基线的最先进的离线RL基线的性能。
We introduce a framework that abstracts Reinforcement Learning (RL) as a sequence modeling problem. This allows us to draw upon the simplicity and scalability of the Transformer architecture, and associated advances in language modeling such as GPT-x and BERT. In particular, we present Decision Transformer, an architecture that casts the problem of RL as conditional sequence modeling. Unlike prior approaches to RL that fit value functions or compute policy gradients, Decision Transformer simply outputs the optimal actions by leveraging a causally masked Transformer. By conditioning an autoregressive model on the desired return (reward), past states, and actions, our Decision Transformer model can generate future actions that achieve the desired return. Despite its simplicity, Decision Transformer matches or exceeds the performance of state-of-the-art model-free offline RL baselines on Atari, OpenAI Gym, and Key-to-Door tasks.