Model-Based Offline Meta-Reinforcement Learning with Regularization

Model-Based Offline Meta-Reinforcement Learning with Regularization
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Sen Lin;Jialin Wan;Tengyu Xu;Yingbin Liang;Junshan Zhang
Sen Lin;Jialin Wan;Tengyu Xu;Yingbin Liang;Junshan Zhang
中科院分区:
其他
文献类型:
--
作者:
Sen Lin;Jialin Wan;Tengyu Xu;Yingbin Liang;Junshan Zhang

文献摘要

相似文献

现有的离线强化学习(RL)方法面临着一些重大挑战,特别是学习策略和行为策略之间的分布转移。离线Meta-RL正在成为解决这些挑战的一种有前途的方法,旨在从一系列任务中学习一种信息性的元策略。然而,正如我们的实证研究所显示的那样,离线Meta-RL在具有良好数据集质量的任务上的表现可能被离线单任务RL方法所超越,这表明必须在通过遵循元策略来“探索”非分布状态动作和通过紧跟行为策略来“利用”离线数据集之间进行微妙的平衡。基于这样的经验分析,我们探索了基于模型的带正则策略优化的离线Meta-RL(MerPO),它学习了一个用于高效任务结构推理的元模型和一个用于安全探测非分布状态动作的信息元策略。特别是,我们设计了一种新的基于元正则化模型的Actor-Critic(RAC)任务内策略优化方法,作为MerPO的关键构建块,使用保守的策略评估和正则化的策略改进,并通过在基于行为策略和基于元策略的两个正则化器之间取得适当的平衡来实现内在的权衡。我们从理论上证明了学习策略在行为策略和元策略上都有保证的改进,从而保证了离线Meta-RL对新任务的性能提升。实验证明,MerPO的性能优于现有的离线Meta-RL方法。
Existing offline reinforcement learning (RL) methods face a few major challenges, particularly the distributional shift between the learned policy and the behavior policy. Offline Meta-RL is emerging as a promising approach to address these challenges, aiming to learn an informative meta-policy from a collection of tasks. Nevertheless, as shown in our empirical studies, offline Meta-RL could be outperformed by offline single-task RL methods on tasks with good quality of datasets, indicating that a right balance has to be delicately calibrated between"exploring"the out-of-distribution state-actions by following the meta-policy and"exploiting"the offline dataset by staying close to the behavior policy. Motivated by such empirical analysis, we explore model-based offline Meta-RL with regularized Policy Optimization (MerPO), which learns a meta-model for efficient task structure inference and an informative meta-policy for safe exploration of out-of-distribution state-actions. In particular, we devise a new meta-Regularized model-based Actor-Critic (RAC) method for within-task policy optimization, as a key building block of MerPO, using conservative policy evaluation and regularized policy improvement; and the intrinsic tradeoff therein is achieved via striking the right balance between two regularizers, one based on the behavior policy and the other on the meta-policy. We theoretically show that the learnt policy offers guaranteed improvement over both the behavior policy and the meta-policy, thus ensuring the performance improvement on new tasks via offline Meta-RL. Experiments corroborate the superior performance of MerPO over existing offline Meta-RL methods.