Detection-averse optimal and receding-horizon control for Markov decision processes

Detection-averse optimal and receding-horizon control for Markov decision processes
复制标题

DOI:
10.1016/j.automatica.2020.109278
复制
发表时间:
2019-08
期刊:
Autom.
影响因子:
--
通讯作者:
Nan I. Li;I. Kolmanovsky;A. Girard
Nan I. Li;I. Kolmanovsky;A. Girard
中科院分区:
其他
文献类型:
--
作者:
Nan I. Li;I. Kolmanovsky;A. Girard

文献摘要

相似文献

在本文中,我们考虑一个马尔可夫决策过程(MDP)中,自我代理打算隐藏其状态检测的对手,同时追求一个名义上的目标。在描述了检测厌恶MDP问题之后,我们首先描述了一种精确求解该问题的值迭代(VI)方法,然后为了克服“维数灾难”,从而获得更大规模问题的可扩展性,我们提出了一种滚动时域优化(RHO)方法来计算近似解。数值例子来说明和比较VI和RHO方法,并显示潜在的实际应用中提出的问题制定。
In this paper, we consider a Markov decision process (MDP) in which the ego agent intends to hide its state from detection by an adversary while pursuing a nominal objective. After formulating the detection-averse MDP problem, we first describe a value iteration (VI) approach to exactly solve it. To overcome the “curse of dimensionality” and thus gain scalability to larger-sized problems, we then propose a receding-horizon optimization (RHO) approach to compute approximate solutions. Numerical examples are reported to illustrate and compare the VI and RHO approaches, and show the potential of the proposed problem formulation for practical applications.