课题基金 / 基金详情

New Computational Approaches for Markov Decision Processes

New Computational Approaches for Markov Decision Processes
马尔可夫决策过程的新计算方法
批准号:
0323220
负责人:
Michael Fu
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-01-01 至 2008-03-31

项目摘要

项目成果

Michael Fu的其他基金

相似基金

相关文献

中文摘要
翻译
大规模马尔可夫决策过程(MDP),也被称为随机动态规划问题,开发实用的计算解决方案的方法,仍然是一个重要和具有挑战性的研究领域。 原则上可以使用MDP建模的许多现代系统的复杂性已经导致不可能明确地枚举转移概率的模型,但是可以容易地生成样本路径,例如,通过随机模拟模型。 该项目研究解决了另外两个不同但关键的问题:如何最好地分配用于生成样本路径的计算预算,以及如何直接(而不是通过值函数近似间接)生成一组强大的良好策略。 特别是,我们提出的方法的主要推力集中在两个不同的范式:有效的基于采样的方法,使用多臂强盗模型和诱导相关性的价值函数估计;和人口为基础的方法,寻找改进的政策,在传统的政策迭代方法,迭代在一个单一的政策。 后一个推力将集中在无限的地平线问题,假设有一个最佳的平稳政策,而前一种方法的目的是有限的地平线问题,其中必须采用向后归纳动态规划。 算法将被开发,然后分析其性能,如收敛速度和性能的理论界限,然后在特定的应用领域进行测试,以调查其实际效用。 具体的问题领域包括美国式金融衍生品的定价;制造系统的能力规划和预防性维护;以及通信网络。
英文摘要
Developing practical computational solution methods for large-scale Markov Decision Processes (MDPs), also known as stochastic dynamic programming problems, remains an important and challenging research area. The complexity of many modern systems that can in principle be modeled using MDPs have resulted in models for which it is not possible to explicitly enumerate the transition probabilities, but for which sample paths can be easily generated, e.g., via a stochastic simulation model. The project research addresses two other distinct but crucial issues that arise: how best to allocate a computational budget that is used to generate sample paths, and how to produce a robust set of good policies directly (rather than indirectly via value function approximations). In particular, the main thrusts of our proposed approaches center on two distinct paradigms: effective sampling-based methodologies using multi-armed bandit models and induced correlation for value function estimation; and population-based approaches for finding improving policies, in contrast to the traditional policy iteration method, which iterates on a single policy. The latter thrust will focus on infinite horizon problems, where there is assumed an optimal stationary policy, whereas the former approaches are intended for finite horizon problems, where backwards induction dynamic programming must be employed. Algorithms will be developed and then analyzed in terms of their properties such as convergence rate and theoretical bounds on performance, followed by testing on specific application areas to investigate their practical utility. Specific problem domains include the pricing of American-style financial derivatives; capacity planning and preventive maintenance in manufacturing systems; and communication networks.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: SCH: Optimal Desensitization Protocol in Support of a Kidney Paired Donation (KPD) System
CAREER: Maintaining volitional effort during electrical stimulation-assisted stroke rehabilitation
  • 批准号:
    1942402
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $55.0万
  • 财政年份:
    2020
  • 负责人:
    Michael Fu
  • 依托单位:
New Approaches for Simulation-Based Optimal Decision Making
New Simulation-Based Approaches to Solving Markov Decision Processes
国内基金
海外基金
Computational Methods for Analyzing Toponome Data