BRAC+: Improved Behavior Regularized Actor Critic for Offline Reinforcement Learning

BRAC+: Improved Behavior Regularized Actor Critic for Offline Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Chi Zhang;S. Kuppannagari;V. Prasanna
Chi Zhang;S. Kuppannagari;V. Prasanna
中科院分区:
其他
文献类型:
--
作者:
Chi Zhang;S. Kuppannagari;V. Prasanna

文献摘要

相似文献

由于经济和安全问题,与环境进行在线交互以收集数据样本用于训练强化学习(RL)代理并不总是可行的。离线强化学习的目标是通过使用先前收集的数据集学习有效的策略来解决这个问题。标准的非策略RL算法容易高估分布外(较少探索)动作的值,因此不适合离线RL。行为正则化将学习到的策略约束在数据集的支持集内,已被提出来解决标准策略外算法的局限性。在本文中,我们改进了行为正则化离线强化学习,并提出了BRAC+。首先,我们提出了量化的分布外的行动,并进行比较使用Kullback-Leibler分歧与使用最大均值离散作为正则化协议。我们提出了一个分析上界的KL分歧的行为正则化,以减少与基于样本的估计方差。其次,我们在数学上表明,即使在温和的假设下使用行为正则化策略更新,学习的Q值也会发散。这导致了Q值的过高估计和学习策略的性能恶化。为了缓解这个问题,我们增加了一个梯度惩罚项的政策评估目标。通过这样做,保证Q值收敛。在具有挑战性的离线RL基准测试中,BRAC+比基线行为正则化方法高出40%~87%,比最先进的方法高出6%。
Online interactions with the environment to collect data samples for training a Reinforcement Learning (RL) agent is not always feasible due to economic and safety concerns. The goal of Offline Reinforcement Learning is to address this problem by learning effective policies using previously collected datasets. Standard off-policy RL algorithms are prone to overestimations of the values of out-of-distribution (less explored) actions and are hence unsuitable for Offline RL. Behavior regularization, which constraints the learned policy within the support set of the dataset, has been proposed to tackle the limitations of standard off-policy algorithms. In this paper, we improve the behavior regularized offline reinforcement learning and propose BRAC+. First, we propose quantification of the out-of-distribution actions and conduct comparisons between using Kullback-Leibler divergence versus using Maximum Mean Discrepancy as the regularization protocol. We propose an analytical upper bound on the KL divergence as the behavior regularizer to reduce variance associated with sample based estimations. Second, we mathematically show that the learned Q values can diverge even using behavior regularized policy update under mild assumptions. This leads to large overestimations of the Q values and performance deterioration of the learned policy. To mitigate this issue, we add a gradient penalty term to the policy evaluation objective. By doing so, the Q values are guaranteed to converge. On challenging offline RL benchmarks, BRAC+ outperforms the baseline behavior regularized approaches by 40%~87% and the state-of-the-art approach by 6%.