Finite State and Action MDPS

Finite State and Action MDPS
复制标题

有限状态和动作MDPS

DOI:
10.1007/978-1-4615-0805-2_2
复制
发表时间:
2003
期刊:
Lecture Notes in Control and Information Sciences
影响因子:
--
通讯作者:
L. Kallenberg
L. Kallenberg
中科院分区:
--
文献类型:
--
作者:
L. Kallenberg

文献摘要

被引文献

相似文献

在本章中,我们研究了有限状态空间和有限动作空间的马尔可夫决策过程。这是自五十年代末以来发展起来的经典理论。我们考虑有限和无限的地平线模型。对于有限时域模型,通常使用总期望报酬的效用函数。对于无限水平线,效用函数不那么明显。我们考虑几个标准:总折现期望报酬、平均期望报酬和更敏感的最优性准则,包括Blackwell最优性准则。最后,我们将讨论其他一些问题。
In this chapter we study Markov decision processes (MDPs) with finite state and action spaces. This is the classical theory developed since the end of the fifties. We consider finite and infinite horizon models. For the finite horizon model the utility function of the total expected reward is commonly used. For the infinite horizon the utility function is less obvious. We consider several criteria: total discounted expected reward, average expected reward and more sensitive optimality criteria including the Blackwell optimality criterion. We end with a variety of other subjects.