Myopic Solutions of Markov Decision Processes and Stochastic Games

Myopic Solutions of Markov Decision Processes and Stochastic Games
复制标题

马尔可夫决策过程和随机博弈的短视解

DOI:
--
复制
发表时间:
1981
影响因子:
2.7
通讯作者:
M. J. Sobel
M. J. Sobel
中科院分区:
管理学4区
文献类型:
--
作者:
M. J. Sobel

文献摘要

被引文献

相似文献

给出了马尔可夫决策过程具有近视最优和随机对策具有近视平衡点的充分条件。如果一个最优点或平衡点可以从静态优化问题或静态纳什博弈的最优点或平衡点中推导出来,那么这个最优点或平衡点就是“短视的”。主要条件是,每个周期的奖励是当前状态和行为的总和,b每个转移概率取决于所采取的行动,而不取决于发生转移的状态,c一个适当的静态最优点或平衡点是无限可重复的。这些条件被几个动态寡头垄断模型和许多马尔可夫决策过程所满足。
Sufficient conditions are presented for a Markov decision process to have a myopic optimum and for a stochastic game to possess a myopic equilibrium point. An optimum or an equilibrium point is said to be "myopic" if it can be deduced from an optimum or an equilibrium point of a static optimization problem or a static [Nash] game. The principal conditions are a each single period reward is the sum of terms due to the current state and action, b each transition probability depends on the action taken but not on the state from which the transition occurs, and c an appropriate static optimum or equilibrium point is ad infinitum repeatable. These conditions are satisfied by several dynamic oligopoly models and numerous Markov decision processes.