课题基金 / 基金详情

Mathematical Sciences: Topics in Markov Decision Processes

Mathematical Sciences: Topics in Markov Decision Processes
数学科学:马尔可夫决策过程主题
批准号:
9404177
负责人:
Alexander Yushkevich
金额:
$4.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1994
资助国家:
美国
项目状态:
已结题
起止时间:
1994-10-15 至 1996-09-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
9404177马尔可夫决策过程理论是运筹学、管理科学和连续时间随机控制问题数值求解的重要工具。MDP中的敏感标准允许在初始阶段很重要的情况下对单位时间平均奖励标准的选择性不足进行调整,这些标准还提供了对最优化条件和方程的性质的更深层次的洞察。到目前为止,只研究了状态空间有限或可数的MDP的敏感准则。具有连续状态空间的模型在许多应用中更相关,例如,当存在不完整的观测时。该研究的目的是将敏感准则理论推广到具有Borel状态空间的MDP。首先研究具有绝对连续转移函数的模型。对于这类模型,我们期望(通过一般状态空间马氏链的极限定理、预解器的Laurent展开,特别是反馈控制集的聚集和紧致的新技术)获得以下结果:1)词典最优性方程的有效性,2)确定性平稳敏感最优策略的存在性,以及3)获得此类策略的词典策略改进算法的有效性。马尔可夫决策过程为许多控制和管理优化问题的分析提供了重要的工具,其中随机性起着重要作用。它们通常用于在相互竞争的需求之间优化操作的资源分配(例如,维护一定数量的系统备件的库存成本与系统的停机时间成本和在给定时间内修复系统有缺陷部件的成本)。在许多这样的问题中,长期平均成本(或回报)是要优化的自然目标函数。然而,长期平均标准对与初始阶段相关的成本或回报很敏感,在某些情况下,这可能会导致非最优政策。只有当系统的可能状态是有限的或可数的情况下,才成功地分析了敏感准则。然而,当适当的状态空间是一个连续体时,存在许多现实情况,例如,当涉及不完全观测时,通常是这种情况。这项研究的目的是将敏感准则理论推广到具有Borel状态空间的马尔可夫决策过程(Borel状态空间足够普遍,以包括所有已知的重要情况)。
英文摘要
9404177 Yushkevich The Theory of Markov Decision Processes (MDP's) is an important tool in operations research, management sciences, and numerical solution of continuous time stochastic control problems. Sensitive criteria in MDP's permit adjustments for the underselectiveness of the average per unit time reward criterion in cases where initial stages are important and they also provide deeper insight into the nature of optimality conditions and equations. Up to now, sensitive criteria have been studied only for MDP's with finite or countable state spaces. Models with a continuous state space are more relevant in many applications, for instance, when there are incomplete observations. The purpose of the proposed research is to extend the theory of sensitive criteria to MDP's with a Borel state space. First of all models with an absolutely continuous transition function will be studied. For such models, we expect to obtain (by means of limit theorems for general state space Markov chains, Laurent expansions of resolvents, and, especially, new techniques for aggregation and compactification of the set of feedback controls) the following results: 1)validity of the lexicographical optimality equation, 2) existence of deterministic stationary sensitively optimal policies, and 3) effectiveness of a lexicographical policy improvement algorithm to get such a policy. Markov Decision Processes provide important tools for the analysis of many control and management optimization problems in which randomness plays a significant role. They are often used to optimize an operation's resource allocations among competing requirements (for example, inventory costs in maintaining a certain number of spare parts for a system versus down time costs for the system versus costs of repairing defective parts of the system in a given amount of time). In many of these problems, the long-run average cost (or reward) is the natural objective function to optimize. However, long-run average criteria are in sensitive to costs or rewards associated with initial stages, which, under certain circumstances, can lead to nonoptimal policies. Sensitive criteria have only been successfully analyzed for the case when the possible states of the system are finite or countable. There are many realistic situations, however, when the appropriate state space is a continuum, for example, when incomplete observations are involved, which is often the case. It is the purpose of this research to extend the theory of sensitive criteria to Markov Decision Processes with Borel state spaces (which are general enough to include all known situations of importance).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Handbook of the Mathematics of the Arts and Sciences的中文翻译
  • 批准号:
    12226504
  • 项目类别:
    数学天元基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2022
  • 负责人:
    黄朝凌
  • 依托单位:
SCIENCE CHINA: Earth Sciences
Journal of Environmental Sciences
SCIENCE CHINA Information Sciences