Exploratory HJB Equations and Their Convergence

Exploratory HJB Equations and Their Convergence
复制标题

DOI:
10.1137/21m1448185
复制
发表时间:
2021-09
期刊:
SIAM J. Control. Optim.
影响因子:
--
通讯作者:
Wenpin Tang;Y. Zhang;X. Zhou
Wenpin Tang;Y. Zhang;X. Zhou
中科院分区:
其他
文献类型:
--
作者:
Wenpin Tang;Y. Zhang;X. Zhou

文献摘要

相似文献

本文研究了由Wang,Zariphopoulou和Zhou(J.Mach.学习.结果:2020年11月21日)在连续时间和空间的强化学习背景下。我们建立了方程粘性解的适定性和正则性,以及当探索水平衰减到零时,探索性控制问题收敛到经典随机控制问题。然后,我们将一般结果应用于高,徐和周(arXiv:2005.04057,2020)介绍的探索性温度控制问题,以在非凸优化的背景下设计模拟退火(SA)的内生温度计划。我们推导出一个明确的收敛速度为这个问题的探索减少到零,并发现,稳定状态的最佳控制过程的存在,但既不是狄拉克质量的全球最佳,也不是吉布斯措施。
We study the exploratory Hamilton--Jacobi--Bellman (HJB) equation arising from the entropy-regularized exploratory control problem, which was formulated by Wang, Zariphopoulou and Zhou (J. Mach. Learn. Res., 21, 2020) in the context of reinforcement learning in continuous time and space. We establish the well-posedness and regularity of the viscosity solution to the equation, as well as the convergence of the exploratory control problem to the classical stochastic control problem when the level of exploration decays to zero. We then apply the general results to the exploratory temperature control problem, which was introduced by Gao, Xu and Zhou (arXiv:2005.04057, 2020) to design an endogenous temperature schedule for simulated annealing (SA) in the context of non-convex optimization. We derive an explicit rate of convergence for this problem as exploration diminishes to zero, and find that the steady state of the optimally controlled process exists, which is however neither a Dirac mass on the global optimum nor a Gibbs measure.