Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance Guarantees

Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance Guarantees
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
--
影响因子:
--
通讯作者:
K. Badrinath;D. Kalathil
K. Badrinath;D. Kalathil
中科院分区:
其他
文献类型:
--
作者:
K. Badrinath;D. Kalathil

文献摘要

被引文献

相似文献

本文针对具有大状态空间的鲁棒马尔可夫决策过程(RMDP)的无模型强化学习问题进行了研究。RMDP框架的目标是找到一种对由于模拟器模型与现实世界设置不匹配而导致的参数不确定性具有鲁棒性的策略。我们首先提出了鲁棒最小二乘策略评估算法,这是一种用于策略评估的多步在线无模型学习算法。我们使用随机逼近技术证明了该算法的收敛性。然后,我们提出了鲁棒最小二乘策略迭代(RLSPI)算法用于学习最优鲁棒策略。我们还给出了所得策略误差(接近最优性)的一般加权欧几里得范数界。最后,我们在OpenAI Gym的一些基准问题上展示了我们的RLSPI算法的性能。
This paper addresses the problem of model-free reinforcement learning for Robust Markov Decision Process (RMDP) with large state spaces. The goal of the RMDPs framework is to find a policy that is robust against the parameter uncertainties due to the mismatch between the simulator model and real-world settings. We first propose Robust Least Squares Policy Evaluation algorithm, which is a multi-step online model-free learning algorithm for policy evaluation. We prove the convergence of this algorithm using stochastic approximation techniques. We then propose Robust Least Squares Policy Iteration (RLSPI) algorithm for learning the optimal robust policy. We also give a general weighted Euclidean norm bound on the error (closeness to optimality) of the resulting policy. Finally, we demonstrate the performance of our RLSPI algorithm on some benchmark problems from OpenAI Gym.