课题基金 / 基金详情

Algorithms for adaptive near-optimal control

Algorithms for adaptive near-optimal control
自适应近最优控制算法
批准号:
391349-2010
负责人:
Tweed, Douglas
金额:
$1.89万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2011
资助国家:
加拿大
项目状态:
已结题
起止时间:
2011-01-01 至 2012-12-31

项目摘要

项目成果

Tweed, Douglas的其他基金

相似基金

相关文献

中文摘要
翻译
在几乎每一个科学或工程领域,都有我们想要以最佳方式控制的过程,例如最大化速度或精度或燃油效率。但最优控制在计算上要求如此之高,除了简单的任务外,它是遥不可及的。第二个最好的事情可能是接近最优的控制,我们计算出一系列越来越好的控制器,越来越接近最优的控制器。为此,有许多算法,但最有效和最通用的可能是广义哈密顿-雅可比-贝尔曼(GHJB)方程的方法。在这里,我证明了这种方法在重要意义上是间接的,可以通过使用更直接的监督学习形式来改进。GHJB方法是基于这样一个事实:如果我们有一个反馈控制器,并且我们学习计算其成本函数的梯度Grad-J,那么我们可以使用该梯度来定义更好的控制器。然后,我们可以使用新控制器的GRAD-J来定义更好的控制器,依此类推。但GHJB的工作原理是间接的,它不学习Grad-J的最佳近似,而是学习相关函数,并由此推断Grad-J的次优估计。我展示了如何直接学习梯度;例如,我们需要报告受控过程不同状态x的Grad-J(X)的信号,并且我展示了如何使用类似于Euler-Lagrange方程的公式来获得它们。我在最近的控制论文中的测试问题上将这种直接算法与GHJB进行了比较,我表明直接方法产生的控制器更接近最优和更简单,(在一个复杂的任务中)需要的函数求值和可调参数减少了10倍。但还需要更多的测试,还有大量的工作要做,以改进和扩展这种方法。
英文摘要
In almost every field of science or engineering there are processes we would like to control in an optimal way, e.g. maximizing speed or accuracy or fuel efficiency. But optimal control is computationally so demanding that it is out of reach except for simple tasks. The next best thing may be near-optimal control, where we compute a sequence of better and better controllers, moving ever closer to the optimal one. There are many algorithms for this purpose, but the most efficient and versatile is probably the method of generalized Hamilton-Jacobi-Bellman (GHJB) equations. Here I show that this method is in an important sense indirect, and can be improved by using a more direct form of supervised learning. The GHJB method is based on the fact that if we have a feedback controller, and we learn to compute the gradient grad-J of its cost-to-go function, then we can use that gradient to define a better controller. We can then use the new controller's grad-J to define a still-better controller, and so on. But GHJB works indirectly in the sense that it doesn't learn the best approximation to grad-J but instead learns a related function and from that infers a suboptimal estimate of grad-J. I show how it is possible to learn the gradient directly; e.g. we need signals that report grad-J(x) for different states x of the controlled process, and I show how to obtain them using a formula similar to the Euler-Lagrange equation. I compare this direct algorithm with GHJB on test problems from recent control papers, and I show that the direct method yields controllers that are more nearly optimal and simpler, requiring (on one complex task) 10 times fewer function evaluations and adjustable parameters. But much more testing is needed, and there is a great deal of work to be done improving and extending this approach.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Algorithms for adaptive near-optimal control
  • 批准号:
    391349-2010
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2013
  • 负责人:
    Tweed, Douglas
  • 依托单位:
Algorithms for adaptive near-optimal control
  • 批准号:
    391349-2010
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2012
  • 负责人:
    Tweed, Douglas
  • 依托单位:
Algorithms for adaptive near-optimal control
  • 批准号:
    391349-2010
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2010
  • 负责人:
    Tweed, Douglas
  • 依托单位:
国内基金
海外基金
下一代无线通信系统自适应调制技术及跨层设计研究
  • 批准号:
    60802033
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    16.0万元
  • 批准年份:
    2008
  • 负责人:
    刘凯明
  • 依托单位:
由蝙蝠耳轮和鼻叶推导新型仿生自适应波束模型的研究
  • 批准号:
    10774092
  • 项目类别:
    面上项目
  • 资助金额:
    39.0万元
  • 批准年份:
    2007
  • 负责人:
    Rolf Mueller
  • 依托单位: