Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization

Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization
复制标题

DOI:
10.1287/opre.2021.2151
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Shicong Cen;Chen Cheng;Yuxin Chen;Yuting Wei;Yuejie Chi
Shicong Cen;Chen Cheng;Yuxin Chen;Yuting Wei;Yuejie Chi
中科院分区:
其他
文献类型:
--
作者:
Shicong Cen;Chen Cheng;Yuxin Chen;Yuting Wei;Yuejie Chi

文献摘要

相似文献

预处理和正则化使更快的加强学习自然政策梯度(NPG)方法以及受到鼓励探索的熵是最受欢迎的政策优化算法之一,尽管经过经验性的成功,但对NPG方法的理论基础仍然存在。有限。假设访问确切的策略评估,作者表明,算法以惊人的速度与国家行动空间的维度无关。这种融合结果适应广泛的学习率,凸显了预处理和调节在实现快速收敛中的作用。
Preconditioning and Regularization Enable Faster Reinforcement Learning Natural policy gradient (NPG) methods, in conjunction with entropy regularization to encourage exploration, are among the most popular policy optimization algorithms in contemporary reinforcement learning. Despite the empirical success, the theoretical underpinnings for NPG methods remain severely limited. In “Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization”, Cen, Cheng, Chen, Wei, and Chi develop nonasymptotic convergence guarantees for entropy-regularized NPG methods under softmax parameterization, focusing on tabular discounted Markov decision processes. Assuming access to exact policy evaluation, the authors demonstrate that the algorithm converges linearly at an astonishing rate that is independent of the dimension of the state-action space. Moreover, the algorithm is provably stable vis-à-vis inexactness of policy evaluation. Accommodating a wide range of learning rates, this convergence result highlights the role of preconditioning and regularization in enabling fast convergence.