Efficient Policy Learning for Non-Stationary MDPs under Adversarial Manipulation

Efficient Policy Learning for Non-Stationary MDPs under Adversarial Manipulation
复制标题

对抗性操纵下非平稳 MDP 的有效政策学习

DOI:
--
复制
发表时间:
2019
期刊:
arXiv.org
影响因子:
--
通讯作者:
S. Sra
S. Sra
中科院分区:
--
文献类型:
--
作者:
Tiancheng Yu;S. Sra

文献摘要

参考文献

被引文献

相似文献

马尔可夫决策过程(MDP)是一种流行的强化学习模型。然而,它通常使用的静态动力学和回报的假设过于严格,在对抗性、非静态或多智能体问题中不成立。我们研究了一种情节设置,其中MDP的参数在不同的情节中可能不同。我们通过开发一种对抗性强化学习(ARL)算法来学习这种潜在的对抗性MDP的可靠策略,该算法将我们的MDP归结为一系列的emph{对抗性}强盗问题。在现有的无模型方法中,ARL得到了$O(SQRT{SATH^3})$RELARY,它相对于$S$、$A$和$T$是最优的,它对$H$的依赖是最好的(即使对于通常的静态MDP)。
A Markov Decision Process (MDP) is a popular model for reinforcement learning. However, its commonly used assumption of stationary dynamics and rewards is too stringent and fails to hold in adversarial, nonstationary, or multi-agent problems. We study an episodic setting where the parameters of an MDP can differ across episodes. We learn a reliable policy of this potentially adversarial MDP by developing an Adversarial Reinforcement Learning (ARL) algorithm that reduces our MDP to a sequence of emph{adversarial} bandit problems. ARL achieves $O(sqrt{SATH^3})$ regret, which is optimal with respect to $S$, $A$, and $T$, and its dependence on $H$ is the best (even for the usual stationary MDP) among existing model-free methods.
DOI: 10.1007/978-3-030-01554-1_11
发表时间: 2018-08
期刊: --
影响因子: --
作者:
Yuzhe Ma-;Kwang-Sung Jun;Lihong Li;Xiaojin Zhu
通讯作者: Yuzhe Ma-;Kwang-Sung Jun;Lihong Li;Xiaojin Zhu