ON GRADUAL-IMPULSE CONTROL OF CONTINUOUS-TIME MARKOV DECISION PROCESSES WITH EXPONENTIAL UTILITY

ON GRADUAL-IMPULSE CONTROL OF CONTINUOUS-TIME MARKOV DECISION PROCESSES WITH EXPONENTIAL UTILITY
复制标题

DOI:
10.1017/apr.2020.64
复制
发表时间:
2021-06-01
影响因子:
1.2
通讯作者:
Zhang, Yi
Zhang, Yi
中科院分区:
数学4区
文献类型:
--
作者:
Guo, Xin;Kurushima, Aiko;Zhang, Yi

文献摘要

被引文献

相似文献

考虑连续时间马尔可夫决策过程的渐进脉冲控制问题,其中系统性能由总成本的指数效用期望来度量。我们表明,在自然条件下的系统原语,存在一个确定性的固定的最优政策,更一般的一类政策,允许多个同时的脉冲,随机选择的随机影响的脉冲,和累积的跳跃。在用最优性方程刻画了价值函数后,我们将渐进脉冲控制问题归结为一个等价的简单离散时间马尔可夫决策过程,其作用空间是渐进和脉冲作用集的并集.
We consider a gradual-impulse control problem of continuous-time Markov decision processes, where the system performance is measured by the expectation of the exponential utility of the total cost. We show, under natural conditions on the system primitives, the existence of a deterministic stationary optimal policy out of a more general class of policies that allow multiple simultaneous impulses, randomized selection of impulses with random effects, and accumulation of jumps. After characterizing the value function using the optimality equation, we reduce the gradual-impulse control problem to an equivalent simple discrete-time Markov decision process, whose action space is the union of the sets of gradual and impulsive actions.