Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2208.06193
复制
发表时间:
2022-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhendong Wang;Jonathan J. Hunt;Mingyuan Zhou
Zhendong Wang;Jonathan J. Hunt;Mingyuan Zhou
中科院分区:
其他
文献类型:
--
作者:
Zhendong Wang;Jonathan J. Hunt;Mingyuan Zhou

文献摘要

相似文献

离线强化学习是强化学习的一个重要范例,其目的是利用先前收集的静态数据集学习最优策略。标准RL方法通常在这种情况下表现不佳,这是由于分布外动作的函数近似误差。虽然已经提出了各种正则化方法来缓解这个问题,但它们通常受到具有有限表达能力的策略类的约束,这可能导致高度次优的解决方案。在本文中,我们建议将策略表示为扩散模型,这是最近一类高度表达的深度生成模型。我们引入了扩散Q学习(Diffusion-QL),它利用条件扩散模型来表示策略。在我们的方法中,我们学习了一个动作值函数,并在条件扩散模型的训练损失中添加了一个最大化动作值的项,这导致了一个寻求接近行为策略的最佳动作的损失。我们展示了基于扩散模型的策略的表现力,以及扩散模型下行为克隆和策略改进的耦合都有助于Diffusion-QL的出色性能。我们说明了我们的方法相比,以前的作品在一个简单的2D土匪的例子与多模态行为政策的优越性。然后,我们证明了我们的方法可以在大多数D4 RL基准测试任务上实现最先进的性能。
Offline reinforcement learning (RL), which aims to learn an optimal policy using a previously collected static dataset, is an important paradigm of RL. Standard RL methods often perform poorly in this regime due to the function approximation errors on out-of-distribution actions. While a variety of regularization methods have been proposed to mitigate this issue, they are often constrained by policy classes with limited expressiveness that can lead to highly suboptimal solutions. In this paper, we propose representing the policy as a diffusion model, a recent class of highly-expressive deep generative models. We introduce Diffusion Q-learning (Diffusion-QL) that utilizes a conditional diffusion model to represent the policy. In our approach, we learn an action-value function and we add a term maximizing action-values into the training loss of the conditional diffusion model, which results in a loss that seeks optimal actions that are near the behavior policy. We show the expressiveness of the diffusion model-based policy, and the coupling of the behavior cloning and policy improvement under the diffusion model both contribute to the outstanding performance of Diffusion-QL. We illustrate the superiority of our method compared to prior works in a simple 2D bandit example with a multimodal behavior policy. We then show that our method can achieve state-of-the-art performance on the majority of the D4RL benchmark tasks.