课题基金 / 基金详情

Research on Adaptive Estimation and Control of Dynamical Systems

Research on Adaptive Estimation and Control of Dynamical Systems
动力系统自适应估计与控制研究
批准号:
9703812
负责人:
Michael Katehakis
金额:
$10.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-08-01 至 2000-07-31

项目摘要

项目成果

Michael Katehakis的其他基金

相似基金

相关文献

中文摘要
翻译
DMS 9703812自适应估计与控制的研究 动力系统。 迈克尔·N卡特哈克斯和赫伯特·罗宾斯 罗格斯大学 摘要 本研究涉及动态系统的自适应控制。 基本的动态模型被称为“马尔可夫决策过程, 不完全信息”(MDP)问题,其中过渡法 和/或预期的单期奖励可能取决于未知的 参数 这一领域最显著的成果是基于思想 利用分离原理和相关的 确定性等价规则,或一致有效规则 一种称为多臂强盗的顺序分配模型 (MAB)问题. 确定性等价规则的局限性是: i)没有关于收敛速度的要求,以及 (2)在某些情况下,以正概率, 规则可能过早地收敛到错误的参数值, 它最终只使用非最优策略。在后面的研究中,典型的方法是将较大的MDP模型拟合到较小的MAB模型中,将每个确定性策略视为一个 奖励人口(土匪)。 其结果是 由此产生的统计上有效的程序包括 从所有确定性策略中进行采样,否则不进行采样 利用问题的最优化方面。因此,他们成为 由于数据收集的复杂性,范围受到限制。 原因是 在实践中,MDP模型的状态空间往往非常大,并且确定性策略集 巨大的 在最近的工作中,研究人员获得了自适应的 具有数据收集要求的程序, 在最小的情况下, 不可约条件建议研究的一个主要方向 涉及为重要的更普遍的问题制定解决方案, 问题,如i)多链MDP,ii)在有 a)边约束,以及iii)奖励的折扣流。 一 第二个重要目标是开发新的适应性 具有实际实用价值的统计方法 和最优性性质的相关问题的检测 总误差和变化点。 自适应控制的主要思想是计算策略(策略, 或控制规则),用于系统的操作, 系统的未知参数,并在这样做时收敛到一个 对于未知的真实值是最优的策略 参数 应用出现在现代工程的许多领域, 财务和运筹学,如可靠性,维护,质量控制,调度,库存和生产计划。 因此,这种类型的问题已被广泛研究, 文学然而,有效的程序应考虑到并 优化收敛速度是最近才得到的 对于特定的模型,数据收集通常是禁止的 复杂性 拟议研究的一个主要目标是 开发相对简单的自适应控制程序, 合理的计算和内存要求, 实现,对于广泛的一类问题,利用的想法,从 调查人员最近的工作。另一个重要目标是 为在这些领域有用的具体模型开发新方法 软件可靠性(错误检测)和质量控制 (改变点)。 这项研究涉及以下国家关注的战略领域:高性能计算,通信和制造业。
英文摘要
DMS 9703812 Research on Adaptive Estimation and Control of Dynamical Systems. Michael N. Katehakis and Herbert Robbins Rutgers University Abstract This research involves work on adaptive control of dynamic systems. The basic dynamic model is known as the "Markov decision process with incomplete information" (MDP) problem, where the transition law and/or the expected one-period rewards may depend on unknown parameters. The most notable results in this area are based on ideas utilizing either a separation principle and the related certainty-equivalence rule, or uniformly efficient rules for the model of sequential allocation known as the multi-armed bandit (MAB) problem. Limitations of the certainty-equivalence rule are: i) there is no claim on the rate of convergence, and ii) there are cases for which, with positive probability, this rule can prematurely converge to a wrong parameter value so that it eventually uses only a non-optimal policy. The typical approach in the latter studies has been to fit the larger MDP model into the smaller MAB one by considering each deterministic policy as a reward-generating population (bandit). A consequence of this is that the resulting statistically efficient procedures involve sampling from all deterministic policies and do not otherwise utilize the optimization aspect of the problem. Thus, they become limited in scope by data collection complexity. The reason is that in practice the state spaces of MDP models tend to be very large and the set of deterministic policies is immense. In recent work the investigators have obtained adaptive procedures with data collection requirements that are proportional to the number of state - action pairs of the MDP, under a minimal irreducibility condition. A major direction of the proposed research involves the development of solutions for important more general problems such as i) multi-chain MDPs, ii) the case in which there a re side constraints, and iii) discounted streams of rewards. A second important goal is the development of new adaptive statistical methods that possess practically useful implementation and optimality properties for the related problems of detection of total error and change points. The main idea of adaptive control is to compute strategies (policies, or control rules) for the operation of a system that estimate the unknown parameters of the system, and in doing so converge to a strategy that is optimal for the true values of the unknown parameters. Applications arise in many areas of modern engineering, finance, and operations research, such as reliability, maintenance, quality control, scheduling, inventory, and production planning. Consequently, this type of problem has been widely studied in the literature. However, effective procedures that take into account and optimize the speed of convergence have been obtained only recently for specific models, often, with prohibitive data collection complexity. A primary objective of the proposed research is the development of relatively simple adaptive control procedures with reasonable computational and memory requirements for on-line implementation, for a wide class of problems, utilizing ideas from recent work of the investigators. Another important goal is the development of new methods for specific models useful in such areas as software reliability (error detection) and quality control (change points). This research relates to the following strategic areas of national concern: high performance computing, communications, and manufacturing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Theoretical and Algorithmic Advances in Sequential Adaptive Decisions
  • 批准号:
    1662629
  • 项目类别:
    Standard Grant
  • 资助金额:
    $42.83万
  • 财政年份:
    2017
  • 负责人:
    Michael Katehakis
  • 依托单位:
EAGER: Event-Driven, Goal-Oriented Dynamic Resource Deployment
  • 批准号:
    1450743
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2014
  • 负责人:
    Michael Katehakis
  • 依托单位:
Research on Adaptive Sampling and Stochastic Scheduling
  • 批准号:
    8507671
  • 项目类别:
    Standard Grant
  • 资助金额:
    $3.5万
  • 财政年份:
    1985
  • 负责人:
    Michael Katehakis
  • 依托单位:
海外基金