课题基金 / 基金详情

Mathematical Sciences: Topics in Markov Decision Processes

Mathematical Sciences: Topics in Markov Decision Processes
数学科学:马尔可夫决策过程主题
批准号:
9404177
负责人:
Alexander Yushkevich
金额:
$4.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1994
资助国家:
美国
项目状态:
已结题
起止时间:
1994-10-15 至 1996-09-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
马尔可夫决策过程理论(MDP)是运筹学、管理科学和连续时间随机控制问题数值求解中的一个重要工具。MDP中的敏感标准允许在初始阶段很重要的情况下,对单位时间平均奖励标准的选择性不足进行调整,它们还提供了对最优条件和方程性质的更深入了解。到目前为止,仅对状态空间有限或可数的MDP进行了敏感判据的研究。具有连续状态空间的模型在许多应用中更为相关,例如,当存在不完整的观测时。本研究的目的是将敏感准则理论扩展到具有Borel状态空间的MDP。首先研究具有绝对连续过渡函数的模型。对于这样的模型,我们期望获得(通过一般状态空间马尔可夫链的极限定理,解决方案的劳伦特展开,特别是反馈控制集合的聚合和紧化的新技术)以下结果:1)字典最优性方程的有效性,2)确定性平稳敏感最优策略的存在性,以及3)字典策略改进算法的有效性来获得这样的策略。马尔可夫决策过程为分析随机性扮演重要角色的许多控制和管理优化问题提供了重要工具。它们通常用于在竞争需求中优化操作的资源分配(例如,维护系统一定数量的备件的库存成本与系统的停机时间成本相比,与在给定时间内修复系统缺陷部件的成本相比)。在许多此类问题中,长期平均成本(或回报)是需要优化的自然目标函数。然而,长期平均标准对与初始阶段相关的成本或回报很敏感,在某些情况下,这可能导致非最优政策。敏感准则只在系统可能状态有限或可数的情况下才被成功地分析。然而,在许多实际情况下,当适当的状态空间是一个连续体时,例如,当涉及到不完整的观察时,通常就是这种情况。本研究的目的是将敏感准则理论扩展到具有Borel状态空间的马尔可夫决策过程(它足够普遍,可以包含所有已知的重要情况)。
英文摘要
9404177 Yushkevich The Theory of Markov Decision Processes (MDP's) is an important tool in operations research, management sciences, and numerical solution of continuous time stochastic control problems. Sensitive criteria in MDP's permit adjustments for the underselectiveness of the average per unit time reward criterion in cases where initial stages are important and they also provide deeper insight into the nature of optimality conditions and equations. Up to now, sensitive criteria have been studied only for MDP's with finite or countable state spaces. Models with a continuous state space are more relevant in many applications, for instance, when there are incomplete observations. The purpose of the proposed research is to extend the theory of sensitive criteria to MDP's with a Borel state space. First of all models with an absolutely continuous transition function will be studied. For such models, we expect to obtain (by means of limit theorems for general state space Markov chains, Laurent expansions of resolvents, and, especially, new techniques for aggregation and compactification of the set of feedback controls) the following results: 1)validity of the lexicographical optimality equation, 2) existence of deterministic stationary sensitively optimal policies, and 3) effectiveness of a lexicographical policy improvement algorithm to get such a policy. Markov Decision Processes provide important tools for the analysis of many control and management optimization problems in which randomness plays a significant role. They are often used to optimize an operation's resource allocations among competing requirements (for example, inventory costs in maintaining a certain number of spare parts for a system versus down time costs for the system versus costs of repairing defective parts of the system in a given amount of time). In many of these problems, the long-run average cost (or reward) is the natural objective function to optimize. However, long-run average criteria are in sensitive to costs or rewards associated with initial stages, which, under certain circumstances, can lead to nonoptimal policies. Sensitive criteria have only been successfully analyzed for the case when the possible states of the system are finite or countable. There are many realistic situations, however, when the appropriate state space is a continuum, for example, when incomplete observations are involved, which is often the case. It is the purpose of this research to extend the theory of sensitive criteria to Markov Decision Processes with Borel state spaces (which are general enough to include all known situations of importance).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Handbook of the Mathematics of the Arts and Sciences的中文翻译
  • 批准号:
    12226504
  • 项目类别:
    数学天元基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2022
  • 负责人:
    黄朝凌
  • 依托单位:
SCIENCE CHINA: Earth Sciences
Journal of Environmental Sciences
SCIENCE CHINA Information Sciences