A stochastic maximum principle approach for reinforcement learning with parameterized environment

A stochastic maximum principle approach for reinforcement learning with parameterized environment
复制标题

DOI:
10.1016/j.jcp.2023.112238
复制
发表时间:
2022-08
期刊:
J. Comput. Phys.
影响因子:
--
通讯作者:
Richard Archibald;F. Bao;J. Yong
Richard Archibald;F. Bao;J. Yong
中科院分区:
其他
文献类型:
--
作者:
Richard Archibald;F. Bao;J. Yong

文献摘要

相似文献

在这项工作中,我们引入了一个随机最大值原理(SMP)的方法来解决强化学习问题的假设,在环境中的未知数可以参数化的基础上物理知识。对于数值算法的发展,我们采用了一种有效的在线参数估计方法作为我们的探索技术,估计在训练过程中的环境参数,并开发最优策略是通过一个有效的后向行动学习方法SMP框架下的政策改进。数值实验表明,SMP方法的强化学习可以产生可靠的控制策略,和SMP求解器中的梯度下降型优化需要更少的训练集相比,基于标准动态规划原理的方法。
In this work, we introduce a stochastic maximum principle (SMP) approach for solving the reinforcement learning problem with the assumption that the unknowns in the environment can be parameterized based on physics knowledge. For the development of numerical algorithms, we apply an effective online parameter estimation method as our exploration technique to estimate the environment parameter during the training procedure, and the exploitation for the optimal policy is achieved by an efficient backward action learning method for policy improvement under the SMP framework. Numerical experiments are presented to demonstrate that the SMP approach for reinforcement learning can produce reliable control policy, and the gradient descent type optimization in the SMP solver requires less training episodes compared with the standard dynamic programming principle based methods.