Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning

Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning
复制标题

DOI:
10.48550/arxiv.2206.10442
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Haoqi Yuan;Zongqing Lu
Haoqi Yuan;Zongqing Lu
中科院分区:
其他
文献类型:
--
作者:
Haoqi Yuan;Zongqing Lu

文献摘要

相似文献

我们研究离线元强化学习,这是一种实用的强化学习范式,它从离线数据中学习以适应新任务。离线数据的分布由行为策略和任务共同决定。现有的离线元强化学习算法不能区分这些因素,使得任务表示不稳定的行为策略的变化。为了解决这个问题,我们提出了一个任务表示的对比学习框架,该框架对训练和测试中行为策略的分布失配具有鲁棒性。我们设计了一个双层编码器结构,使用互信息最大化来形式化任务表示学习,推导出一个对比学习目标,并介绍了几种方法来近似负对的真实分布。在各种离线元强化学习基准上的实验表明,我们的方法优于现有方法,特别是在推广到分布外行为策略方面。该代码可在https://github.com/PKU-AI-Edge/CORRO上获得。
We study offline meta-reinforcement learning, a practical reinforcement learning paradigm that learns from offline data to adapt to new tasks. The distribution of offline data is determined jointly by the behavior policy and the task. Existing offline meta-reinforcement learning algorithms cannot distinguish these factors, making task representations unstable to the change of behavior policies. To address this problem, we propose a contrastive learning framework for task representations that are robust to the distribution mismatch of behavior policies in training and test. We design a bi-level encoder structure, use mutual information maximization to formalize task representation learning, derive a contrastive learning objective, and introduce several approaches to approximate the true distribution of negative pairs. Experiments on a variety of offline meta-reinforcement learning benchmarks demonstrate the advantages of our method over prior methods, especially on the generalization to out-of-distribution behavior policies. The code is available at https://github.com/PKU-AI-Edge/CORRO.