Online Robust Reinforcement Learning with Model Uncertainty

Online Robust Reinforcement Learning with Model Uncertainty
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Yue Wang;Shaofeng Zou
Yue Wang;Shaofeng Zou
中科院分区:
其他
文献类型:
--
作者:
Yue Wang;Shaofeng Zou

文献摘要

相似文献

鲁棒强化学习(RL)是在不确定的MDP集合上找到一个优化最坏情况性能的策略。在本文中,我们专注于无模型鲁棒RL,其中不确定性集被定义为以错误指定的MDP为中心,该MDP顺序地生成单个样本轨迹,并且假设是未知的。我们开发了一种基于样本的方法来估计未知的不确定性集,并设计了一个强大的Q学习算法(表格的情况下)和强大的TDC算法(函数近似设置),它可以在一个在线和增量的方式来实现。对于鲁棒Q学习算法,我们证明了它收敛到最优鲁棒Q函数,对于鲁棒TDC算法,我们证明了它渐近收敛到某些平稳点。与[Roy等人]中的结果不同,2017],我们的算法不需要任何额外的条件折扣因子,以保证收敛。我们进一步描述了这两种算法的有限时间误差界,并表明鲁棒Q-学习和鲁棒TDC算法收敛速度与其香草同行一样快(在一个常数因子内)。数值实验进一步证明了算法的鲁棒性。我们的方法可以很容易地扩展到鲁棒许多其他算法,例如,TD、SARSA和其他GTD算法。
Robust reinforcement learning (RL) is to find a policy that optimizes the worst-case performance over an uncertainty set of MDPs. In this paper, we focus on model-free robust RL, where the uncertainty set is defined to be centering at a misspecified MDP that generates a single sample trajectory sequentially and is assumed to be unknown. We develop a sample-based approach to estimate the unknown uncertainty set and design a robust Q-learning algorithm (tabular case) and robust TDC algorithm (function approximation setting), which can be implemented in an online and incremental fashion. For the robust Q-learning algorithm, we prove that it converges to the optimal robust Q function, and for the robust TDC algorithm, we prove that it converges asymptotically to some stationary points. Unlike the results in [Roy et al., 2017], our algorithms do not need any additional conditions on the discount factor to guarantee the convergence. We further characterize the finite-time error bounds of the two algorithms and show that both the robust Q-learning and robust TDC algorithms converge as fast as their vanilla counterparts(within a constant factor). Our numerical experiments further demonstrate the robustness of our algorithms. Our approach can be readily extended to robustify many other algorithms, e.g., TD, SARSA, and other GTD algorithms.