Robust Partially Observable Markov Decision Processes

Robust Partially Observable Markov Decision Processes
复制标题

鲁棒的部分可观察马尔可夫决策过程

DOI:
10.2139/ssrn.3195310
复制
发表时间:
2018
期刊:
DecisionSciRN: Simulation Based Optimization (Topic)
影响因子:
--
通讯作者:
S. Saghafian
S. Saghafian
中科院分区:
--
文献类型:
--
作者:
M. Rasouli;S. Saghafian

文献摘要

被引文献

相似文献

在各种应用中,需要在接收到关于底层系统状态的不完美观测之后动态地做出决策。部分可观测马尔可夫决策过程(POMDP)被广泛应用于此类应用。然而,要使用POMDP,决策者必须能够获得每个可能的状态和动作对下的核心状态和观察转移概率的可靠估计。这往往是具有挑战性的,主要是由于缺乏充足的数据,特别是当一些行动在实践中没有足够频繁地采取时。这大大限制了POMDPs在真实的世界环境中的应用。例如,在医疗保健中,医学测试通常会出现假阳性和假阴性错误,因此,决策者对患者的健康状态的信息不完善。此外,由于过去没有推荐或探索某些治疗选项,因此无法使用数据来可靠地估计关于患者健康状态的所有所需转移概率。我们介绍了一个扩展的POMDPs,称为鲁棒POMDPs(RPOMDPs),它允许动态决策时,有歧义的转移概率。这种扩展通过减少对单个概率转换模型的依赖,同时仍然允许不完美的状态观测,从而能够做出鲁棒的决策。我们开发动态规划方程求解RPOMDPs,提供了一个充足的统计和信息状态,讨论了如何降低其计算复杂性,并将它们连接到随机零和游戏与不完美的私人监控。
In a variety of applications, decisions need to be made dynamically after receiving imperfect observations about the state of an underlying system. Partially Observable Markov Decision Processes (POMDPs) are widely used in such applications. To use a POMDP, however, a decision-maker must have access to reliable estimations of core state and observation transition probabilities under each possible state and action pair. This is often challenging mainly due to lack of ample data, especially when some actions are not taken frequently enough in practice. This significantly limits the application of POMDPs in real world settings. In healthcare, for example, medical tests are typically subject to false-positive and false-negative errors, and hence, the decision-maker has imperfect information about the health state of a patient. Furthermore, since some treatment options have not been recommended or explored in the past, data cannot be used to reliably estimate all the required transition probabilities regarding the health state of the patient. We introduce an extension of POMDPs, termed Robust POMDPs (RPOMDPs), which allows dynamic decision-making when there is ambiguity regarding transition probabilities. This extension enables making robust decisions by reducing the reliance on a single probabilistic model of transitions, while still allowing for imperfect state observations. We develop dynamic programming equations for solving RPOMDPs, provide a sucient statistic and an information state, discuss ways in which their computational complexity can be reduced, and connect them to stochastic zero-sum games with imperfect private monitoring.