Model-Free Robust Average-Reward Reinforcement Learning

Model-Free Robust Average-Reward Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2305.10504
复制
发表时间:
2023-05
期刊:
--
影响因子:
--
通讯作者:
Yue Wang;Alvaro Velasquez;George K. Atia;Ashley Prater-Bennette;Shaofeng Zou
Yue Wang;Alvaro Velasquez;George K. Atia;Ashley Prater-Bennette;Shaofeng Zou
中科院分区:
其他
文献类型:
--
作者:
Yue Wang;Alvaro Velasquez;George K. Atia;Ashley Prater-Bennette;Shaofeng Zou

文献摘要

相似文献

鲁棒马尔可夫决策过程(MDP)通过优化不确定MDP集的最坏情况性能来解决模型不确定性的挑战。在本文中,我们专注于鲁棒平均报酬MDP下的无模型设置。我们首先从理论上刻画了鲁棒平均报酬Bellman方程解的结构,这对于我们后面的收敛性分析是必不可少的。然后,我们设计了两个无模型算法,鲁棒相对值迭代(RVI)TD和鲁棒RVI Q学习,并从理论上证明了它们的收敛到最优解。我们提供了几个广泛使用的不确定性集作为例子,包括那些定义的污染模型,总变差,卡方散度,Kullback-Leibler(KL)散度和Wasserstein距离。
Robust Markov decision processes (MDPs) address the challenge of model uncertainty by optimizing the worst-case performance over an uncertainty set of MDPs. In this paper, we focus on the robust average-reward MDPs under the model-free setting. We first theoretically characterize the structure of solutions to the robust average-reward Bellman equation, which is essential for our later convergence analysis. We then design two model-free algorithms, robust relative value iteration (RVI) TD and robust RVI Q-learning, and theoretically prove their convergence to the optimal solution. We provide several widely used uncertainty sets as examples, including those defined by the contamination model, total variation, Chi-squared divergence, Kullback-Leibler (KL) divergence and Wasserstein distance.