Inverse stochastic optimal controls

Inverse stochastic optimal controls
复制标题

DOI:
10.1016/j.automatica.2022.110831
复制
发表时间:
2020-05
期刊:
Autom.
影响因子:
--
通讯作者:
Yumiharu Nakano
Yumiharu Nakano
中科院分区:
其他
文献类型:
--
作者:
Yumiharu Nakano

文献摘要

相似文献

研究了性能指标为控制过程的二次罚项的一般扩散随机最优控制的反问题。在系统动力学、代价函数和最优控制过程的温和条件下,我们利用随机极大值原理证明了我们的反问题是适定的。然后,利用适定性,我们将反问题归结为值函数所涉及的随机变量的期望的求根问题,该问题有唯一解。基于这一结果,我们提出了一种反问题的数值方法,用观测到的最优控制过程和相应的状态过程的算术平均值来代替上面的期望。哈密顿-雅可比-贝尔曼方程的数值分析的最新进展使所提出的方法能够在多维情况下实现。特别地,在Hamilton-Jacobi-Bellman方程的基于核的配置法的帮助下,我们的方法即使在值函数的显式形式不可用的情况下仍然很好地求解逆问题。数值实验表明,该方法能够较高精度地恢复未知的惩罚参数。
We study an inverse problem of the stochastic optimal control of general diffusions with performance index having the quadratic penalty term of the control process. Under mild conditions on the system dynamics, the cost functions, and the optimal control process, we show that our inverse problem is well-posed using a stochastic maximum principle. Then, with the well-posedness, we reduce the inverse problem to some root finding problem of the expectation of a random variable involved with the value function, which has a unique solution. Based on this result, we propose a numerical method for our inverse problem by replacing the expectation above with arithmetic mean of observed optimal control processes and the corresponding state processes. The recent progress of numerical analysis of Hamilton–Jacobi–Bellman equations enables the proposed method to be implementable for multi-dimensional cases. In particular, with the help of the kernel-based collocation method for Hamilton–Jacobi–Bellman equations, our method for the inverse problems still works well even when an explicit form of the value function is unavailable. Several numerical experiments show that the numerical method recovers the unknown penalty parameter with high accuracy.