Markov Decision Processes: Discrete Stochastic Dynamic Programming

Markov Decision Processes: Discrete Stochastic Dynamic Programming
复制标题

DOI:
10.2307/2291177
复制
发表时间:
1994-04
影响因子:
4.4
通讯作者:
M. Puterman
M. Puterman
中科院分区:
医学2区
文献类型:
--
作者:
M. Puterman

文献摘要

被引文献

相似文献

来自出版商:在过去的十年中,马尔可夫决策过程的理论和应用研究,以及越来越多地使用这些模型在生态学,经济学,通信工程和其他领域的结果是不确定的,顺序决策过程是必要的。对这种增加的活动的及时反应,马丁L。普特曼的新工作提供了一个独特的最新的,统一的,严格的理论,计算和应用研究马尔可夫决策过程模型的治疗。它讨论了该领域的所有主要研究方向,突出了马尔可夫决策过程模型的许多重要应用,并探讨了许多以前被忽视或在文献中粗略报道的重要主题。马尔可夫决策过程主要关注无限时域离散时间模型和离散时间空间模型,同时也研究具有任意状态空间的模型,有限时域模型和连续时间离散状态模型。这本书是围绕最优性标准组织的,使用一个以最优性(贝尔曼)方程为中心的通用框架来呈现结果。结果以“定理证明”的形式呈现,并通过讨论和示例进行详细说明,包括任何其他书籍中没有的结果。第3章中提出的两状态马尔可夫决策过程模型在整本书中被反复分析,并演示了许多结果和算法。马尔可夫决策过程涵盖了最近的研究进展,如可数状态空间模型与平均回报标准,约束模型,模型与风险敏感的最优性标准。它还探讨了几个在其他书籍中很少或根本没有受到关注的主题,包括修改的策略迭代,具有平均奖励标准的多链模型和敏感最优性。此外,每一章都有一个书目注释部分,对相关历史文献进行了评论。
From the Publisher: The past decade has seen considerable theoretical and applied research on Markov decision processes, as well as the growing use of these models in ecology, economics, communications engineering, and other fields where outcomes are uncertain and sequential decision-making processes are needed. A timely response to this increased activity, Martin L. Puterman's new work provides a uniquely up-to-date, unified, and rigorous treatment of the theoretical, computational, and applied research on Markov decision process models. It discusses all major research directions in the field, highlights many significant applications of Markov decision processes models, and explores numerous important topics that have previously been neglected or given cursory coverage in the literature. Markov Decision Processes focuses primarily on infinite horizon discrete time models and models with discrete time spaces while also examining models with arbitrary state spaces, finite horizon models, and continuous-time discrete state models. The book is organized around optimality criteria, using a common framework centered on the optimality (Bellman) equation for presenting results. The results are presented in a "theorem-proof" format and elaborated on through both discussion and examples, including results that are not available in any other book. A two-state Markov decision process model, presented in Chapter 3, is analyzed repeatedly throughout the book and demonstrates many results and algorithms. Markov Decision Processes covers recent research advances in such areas as countable state space models with average reward criterion, constrained models, and models with risk sensitive optimality criteria. It also explores several topics that have received little or no attention in other books, including modified policy iteration, multichain models with average reward criterion, and sensitive optimality. In addition, a Bibliographic Remarks section in each chapter comments on relevant historic