Kullback-Leibler control for discrete-time nonlinear systems on continuous spaces

Kullback-Leibler control for discrete-time nonlinear systems on continuous spaces
复制标题

连续空间上离散时间非线性系统的 Kullback-Leibler 控制

DOI:
10.1080/18824889.2022.2095827
复制
发表时间:
2022
期刊:
SICE Journal of Control, Measurement, and System Integration
影响因子:
--
通讯作者:
Kenji Kashima
Kenji Kashima
中科院分区:
--
文献类型:
--
作者:
Kaito Ito;Kenji Kashima

文献摘要

相似文献

Kullback-Leibler(KL)控制使非线性最优控制问题的有效数值方法成为可能。KL控制的关键假设是过渡分布完全可控。然而,当动力学在连续空间中演化时,这一假设常常被违反。因此,将KL控制应用于具有连续空间的问题需要一些近似,这导致最优性的损失。为了避免这种近似,在本文中,我们重新制定连续空间的KL控制问题,使其不需要不切实际的假设。原KL控制与改进后的KL控制的主要区别在于前者用受控与非受控跃迁分布之间的KL发散度来衡量控制效果,而后者用噪声驱动跃迁代替非受控跃迁。我们表明,重新KL控制承认有效的数值算法,如原来的一个没有不合理的假设。具体地,可以通过使用基于其路径积分表示的蒙特卡罗方法来计算相关联的值函数。
Kullback–Leibler (KL) control enables efficient numerical methods for nonlinear optimal control problems. The crucial assumption of KL control is the full controllability of transition distributions. However, this assumption is often violated when the dynamics evolves in a continuous space. Consequently, applying KL control to problems with continuous spaces requires some approximation, which leads to the loss of the optimality. To avoid such an approximation, in this paper, we reformulate the KL control problem for continuous spaces so that it does not require unrealistic assumptions. The key difference between the original and reformulated KL control is that the former measures the control effort by the KL divergence between controlled anduncontrolledtransition distributions while the latter replaces the uncontrolled transition by anoise-driventransition. We show that the reformulated KL control admits efficient numerical algorithms like the original one without unreasonable assumptions. Specifically, the associated value function can be computed by using a Monte Carlo method based on its path integral representation.