Adversarial Model for Offline Reinforcement Learning

Adversarial Model for Offline Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2302.11048
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
M. Bhardwaj;Tengyang Xie;Byron Boots;Nan Jiang;Ching-An Cheng
M. Bhardwaj;Tengyang Xie;Byron Boots;Nan Jiang;Ching-An Cheng
中科院分区:
其他
文献类型:
--
作者:
M. Bhardwaj;Tengyang Xie;Byron Boots;Nan Jiang;Ching-An Cheng

文献摘要

被引文献

相似文献

我们提出了一种新的基于模型的离线强化学习(RL)框架,称为离线强化学习的对抗模型(ARMOR),它可以鲁棒地学习策略,以改进任意参考策略,而不管数据覆盖率如何。ARMOR旨在通过逆向训练马尔可夫决策过程模型来优化相对于参考策略的最坏情况性能的策略。在理论上,我们证明了ARMOR,与一个良好的调整超参数,可以竞争的最佳政策的数据覆盖范围内的参考政策的数据支持。与此同时,ARMOR对超参数选择是鲁棒的:由ARMOR学习的策略,具有“任何“可接受的超参数,永远不会降低参考策略的性能,即使参考策略没有被数据集覆盖。为了在实践中验证这些属性,我们设计了一个可扩展的ARMOR实现,通过对抗训练,与典型的基于模型的方法相比,它可以在不使用模型集成的情况下优化策略。我们表明,ARMOR实现了最先进的离线无模型和基于模型的RL算法的性能,并可以鲁棒地提高各种超参数选择的参考策略。
We propose a novel model-based offline Reinforcement Learning (RL) framework, called Adversarial Model for Offline Reinforcement Learning (ARMOR), which can robustly learn policies to improve upon an arbitrary reference policy regardless of data coverage. ARMOR is designed to optimize policies for the worst-case performance relative to the reference policy through adversarially training a Markov decision process model. In theory, we prove that ARMOR, with a well-tuned hyperparameter, can compete with the best policy within data coverage when the reference policy is supported by the data. At the same time, ARMOR is robust to hyperparameter choices: the policy learned by ARMOR, with"any"admissible hyperparameter, would never degrade the performance of the reference policy, even when the reference policy is not covered by the dataset. To validate these properties in practice, we design a scalable implementation of ARMOR, which by adversarial training, can optimize policies without using model ensembles in contrast to typical model-based methods. We show that ARMOR achieves competent performance with both state-of-the-art offline model-free and model-based RL algorithms and can robustly improve the reference policy over various hyperparameter choices.