Reinforcement Learning in Robust Markov Decision Processes

Reinforcement Learning in Robust Markov Decision Processes
复制标题

DOI:
10.1287/moor.2016.0779
复制
发表时间:
2016-11-01
影响因子:
1.7
通讯作者:
Mannor, Shie
Mannor, Shie
中科院分区:
数学2区
文献类型:
--
作者:
Lim, Shiau Hong;Xu, Huan;Mannor, Shie

文献摘要

被引文献

相似文献

马尔可夫决策过程(MDP)中的一个重要挑战是确保对意外或对抗性系统行为的鲁棒性。应对这一挑战的标准范式是稳健的MDP框架,该框架将参数建模为预定义“不确定性集”的任意元素,并寻求极大极小策略--在不确定性集中参数的最差实现下表现最佳的策略。强大的MDP框架的一个关键问题,在文献中基本上没有解决,是如何找到一个原则性的数据驱动的方式的不确定性的适当描述。在本文中,我们解决这个问题,使用在线学习的方法:我们设计了一个算法,不知道真正的不确定性模型,能够适应其保护水平的不确定性,并在长期执行以及极大极小化政策,如果真正的不确定性模型是已知的。事实上,该算法实现了与标准MDP相似的遗憾界限,其中没有参数是对抗性的,这表明几乎没有额外的成本,我们可以适应鲁棒学习来处理MDP中的不确定性。据我们所知,这是第一次尝试学习稳健MDPs中的不确定性。
An important challenge in Markov decision processes (MDP) is to ensure robustness with respect to unexpected or adversarial system behavior. A standard paradigm to tackle this challenge is the robust MDP framework that models the parameters as arbitrary elements of pre-defined "uncertainty sets," and seeks the minimax policy-the policy that performs the best under the worst realization of the parameters in the uncertainty set. A crucial issue of the robust MDP framework, largely unaddressed in literature, is how to find appropriate description of the uncertainty in a principled data-driven way. In this paper we address this problem using an online learning approach: we devise an algorithm that, without knowing the true uncertainty model, is able to adapt its level of protection to uncertainty, and in the long run performs as well as the minimax policy as if the true uncertainty model is known. Indeed, the algorithm achieves similar regret bounds as standard MDP where no parameter is adversarial, which shows that with virtually no extra cost we can adapt robust learning to handle uncertainty in MDPs. To the best of our knowledge, this is the first attempt to learn uncertainty in robust MDPs.