Bayesian inference approach for entropy regularized reinforcement learning with stochastic dynamics

Bayesian inference approach for entropy regularized reinforcement learning with stochastic dynamics
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Argenis Arriojas;Jacob Adamczyk;Stas Tiomkin;R. Kulkarni
Argenis Arriojas;Jacob Adamczyk;Stas Tiomkin;R. Kulkarni
中科院分区:
其他
文献类型:
--
作者:
Argenis Arriojas;Jacob Adamczyk;Stas Tiomkin;R. Kulkarni

文献摘要

相似文献

我们开发了一种新的方法来确定熵正则化强化学习(RL)与随机动力学的最优策略。对于确定性动力学,最优策略可以在控制即推理框架中使用贝叶斯推理导出;然而,对于随机动力学,直接使用这种方法会导致冒险的乐观策略。为了解决这个问题,熵正则化RL中的当前方法涉及约束优化过程,该过程将系统动力学固定到原始动力学,但是这种方法与无约束贝叶斯推理框架不一致。在这项工作中,我们解决了这种不一致性,通过开发一个精确的映射,从熵正则化RL的约束优化问题到一个不同的优化问题,可以使用无约束贝叶斯推理方法来解决。我们表明,这两个问题的最优策略是相同的,因此我们的结果导致的熵正则化RL随机动态通过贝叶斯推理的最优策略的精确解。
We develop a novel approach to determine the optimal policy in entropy-regularized reinforcement learning (RL) with stochastic dynamics. For deterministic dynamics, the optimal policy can be derived using Bayesian inference in the control-as-inference framework; however, for stochastic dynamics, the direct use of this approach leads to risk-taking optimistic policies. To address this issue, current approaches in entropy-regularized RL involve a constrained optimization procedure which fixes system dynamics to the original dynamics, however this approach is not consistent with the unconstrained Bayesian inference framework. In this work we resolve this inconsistency by developing an exact mapping from the constrained optimization problem in entropy-regularized RL to a different optimization problem which can be solved using the unconstrained Bayesian inference approach. We show that the optimal policies are the same for both problems, thus our results lead to the exact solution for the optimal policy in entropy-regularized RL with stochastic dynamics through Bayesian inference.