Policy Mirror Descent for Regularized Reinforcement Learning: A Generalized Framework with Linear Convergence

Policy Mirror Descent for Regularized Reinforcement Learning: A Generalized Framework with Linear Convergence
复制标题

DOI:
10.1137/21m1456789
复制
发表时间:
2021-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Wenhao Zhan;Shicong Cen;Baihe Huang;Yuxin Chen;Jason D. Lee;Yuejie Chi
Wenhao Zhan;Shicong Cen;Baihe Huang;Yuxin Chen;Jason D. Lee;Yuejie Chi
中科院分区:
其他
文献类型:
--
作者:
Wenhao Zhan;Shicong Cen;Baihe Huang;Yuxin Chen;Jason D. Lee;Yuejie Chi

文献摘要

相似文献

策略优化是通过优化技术最大化价值函数来找到所需策略的,这是加固学习的核心(RL)。除了价值最大化之外,还出现了其他实际考虑,包括鼓励探索的需求,以及由于安全,资源和运营限制而确保学习政策的某些结构性属性。这些通常可以通过正规化的RL来解释,该RL可以通过结构促进的正规器来增强目标价值函数。为了关注折扣的无限马尔可夫决策过程,我们提出了一种通用政策镜下降(GPMD)算法,用于求解正则化RL。作为策略镜下降的概括(Arxiv:2102.00135),我们的算法可容纳一般的凸正规化器类别,并促进了Bregman Divergence在使用常规化合物中的使用。我们证明,即使正规器缺乏强大的凸度和光滑度,我们的算法在整个学习率的整个学习速率上也将线性收敛到全球解决方案。此外,在面对不精确的策略评估和不完善的策略更新时,这种线性收敛功能是稳定的。提供数值实验以证实GPMD的吸引力。
Policy optimization, which finds the desired policy by maximizing value functions via optimization techniques, lies at the heart of reinforcement learning (RL). In addition to value maximization, other practical considerations arise as well, including the need of encouraging exploration, and that of ensuring certain structural properties of the learned policy due to safety, resource and operational constraints. These can often be accounted for via regularized RL, which augments the target value function with a structure-promoting regularizer. Focusing on discounted infinite-horizon Markov decision processes, we propose a generalized policy mirror descent (GPMD) algorithm for solving regularized RL. As a generalization of policy mirror descent (arXiv:2102.00135), our algorithm accommodates a general class of convex regularizers and promotes the use of Bregman divergence in cognizant of the regularizer in use. We demonstrate that our algorithm converges linearly to the global solution over an entire range of learning rates, in a dimension-free fashion, even when the regularizer lacks strong convexity and smoothness. In addition, this linear convergence feature is provably stable in the face of inexact policy evaluation and imperfect policy updates. Numerical experiments are provided to corroborate the appealing performance of GPMD.