The relationship between the maximum principle and dynamic programming

The relationship between the maximum principle and dynamic programming
复制标题

极大值原理与动态规划的关系

DOI:
10.1137/0325071
复制
发表时间:
1987
影响因子:
2.2
通讯作者:
Richard D. Vinter
Richard D. Vinter
中科院分区:
数学2区
文献类型:
--
作者:
F. Clarke;Richard D. Vinter

文献摘要

被引文献

相似文献

设$V(t,x)$为最优控制问题的最小代价,视为初始时间和状态$(t,x)$的函数。动态规划关注$V( \cdot , \cdot )$的性质,特别是将其表征为Hamilton-Jacobi-Bellman方程的解。关于最大值原则和动态规划的启发式论证已经提出了很长时间,根据\[p(t) = - V_x \left( {t,x_0 (t)} \right).\],其中$x_0 ( \cdot )$是考虑的最小状态函数,$p( \cdot )$是最大值原则的协态函数。在本文中,我们检验了这种说法的有效性,并发现这种关系,被解释为涉及广义梯度的微分包含,对于非常大的一类非光滑最优控制问题,几乎在任何地方和端点上都是正确的。
Let $V(t,x)$ be the infimum cost of an optimal control problem, viewed as a function of the initial time and state $(t,x)$. Dynamic Programming is concerned with the properties of $V( \cdot , \cdot )$ and in particular with its characterization as a solution to the Hamilton–Jacobi–Bellman equation. Heuristic arguments have long been advanced relating the Maximum Principle to Dynamic Programming according to \[p(t) = - V_x \left( {t,x_0 (t)} \right).\] Here $x_0 ( \cdot )$ is the minimizing state function under consideration and $p( \cdot )$ is the costate function of the Maximum Principle. In this paper we examine the validity of such claims and find that this relationship, interpreted as a differential inclusion involving the generalized gradient, is indeed true, almost everywhere and at the endpoints, for a very large class of nonsmooth optimal control problems.