Algorithms for adaptive near-optimal control
Algorithms for adaptive near-optimal control
批准号:
391349-2010
负责人:
Tweed, Douglas
金额:
$1.89万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2013
资助国家:
加拿大
项目状态:
已结题
起止时间:
2013-01-01 至 2014-12-31
中文摘要
在几乎每一个科学或工程领域,都有我们希望以最佳方式控制的过程,例如,最大限度地提高速度、精度或燃油效率。但最优控制在计算上的要求如此之高,以至于除了简单的任务外,它是遥不可及的。下一个最佳选择可能是接近最优控制,我们计算出一系列越来越好的控制器,越来越接近最优。有许多算法用于此目的,但最有效和最通用的可能是广义Hamilton-Jacobi-Bellman (GHJB)方程方法。在这里,我表明这种方法在重要意义上是间接的,可以通过使用更直接的监督学习形式来改进。GHJB方法是基于这样一个事实,如果我们有一个反馈控制器,我们学习计算梯度梯度梯度- j的成本函数,然后我们可以使用这个梯度来定义一个更好的控制器。然后,我们可以使用新控制器的grad-J来定义一个更好的控制器,以此类推。但是GHJB的作用是间接的,因为它没有学习到grad-J的最佳近似值,而是学习了一个相关的函数,并从中推断出grad-J的次优估计。我展示了如何直接学习梯度;例如,我们需要对受控过程的不同状态x报告grad-J(x)的信号,我展示了如何使用类似于欧拉-拉格朗日方程的公式来获得它们。我将这种直接算法与GHJB在最近的控制论文中的测试问题上进行了比较,我表明,直接方法产生的控制器更接近最优,更简单,(在一个复杂的任务上)需要的函数评估和可调参数减少了10倍。但是需要进行更多的测试,并且需要做大量的工作来改进和扩展这种方法。
英文摘要
In almost every field of science or engineering there are processes we would like to control in an optimal way, e.g. maximizing speed or accuracy or fuel efficiency. But optimal control is computationally so demanding that it is out of reach except for simple tasks. The next best thing may be near-optimal control, where we compute a sequence of better and better controllers, moving ever closer to the optimal one. There are many algorithms for this purpose, but the most efficient and versatile is probably the method of generalized Hamilton-Jacobi-Bellman (GHJB) equations. Here I show that this method is in an important sense indirect, and can be improved by using a more direct form of supervised learning. The GHJB method is based on the fact that if we have a feedback controller, and we learn to compute the gradient grad-J of its cost-to-go function, then we can use that gradient to define a better controller. We can then use the new controller's grad-J to define a still-better controller, and so on. But GHJB works indirectly in the sense that it doesn't learn the best approximation to grad-J but instead learns a related function and from that infers a suboptimal estimate of grad-J. I show how it is possible to learn the gradient directly; e.g. we need signals that report grad-J(x) for different states x of the controlled process, and I show how to obtain them using a formula similar to the Euler-Lagrange equation. I compare this direct algorithm with GHJB on test problems from recent control papers, and I show that the direct method yields controllers that are more nearly optimal and simpler, requiring (on one complex task) 10 times fewer function evaluations and adjustable parameters. But much more testing is needed, and there is a great deal of work to be done improving and extending this approach.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Algorithms for adaptive near-optimal control
-
批准号:391349-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.89万
-
财政年份:2012
-
负责人:Tweed, Douglas
-
依托单位:
Algorithms for adaptive near-optimal control
-
批准号:391349-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.89万
-
财政年份:2011
-
负责人:Tweed, Douglas
-
依托单位:
Algorithms for adaptive near-optimal control
-
批准号:391349-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.89万
-
财政年份:2010
-
负责人:Tweed, Douglas
-
依托单位:
国内基金
海外基金
下一代无线通信系统自适应调制技术及跨层设计研究
-
批准号:60802033
-
项目类别:青年科学基金项目
-
资助金额:16.0万元
-
批准年份:2008
-
负责人:刘凯明
-
依托单位:
由蝙蝠耳轮和鼻叶推导新型仿生自适应波束模型的研究
-
批准号:10774092
-
项目类别:面上项目
-
资助金额:39.0万元
-
批准年份:2007
-
负责人:Rolf Mueller
-
依托单位: