Softmax policy gradient methods can take exponential time to converge

Softmax policy gradient methods can take exponential time to converge
复制标题

DOI:
10.1007/s10107-022-01920-6
复制
发表时间:
2021-02
影响因子:
2.7
通讯作者:
Gen Li;Yuting Wei;Yuejie Chi;Yuantao Gu;Yuxin Chen
Gen Li;Yuting Wei;Yuejie Chi;Yuantao Gu;Yuxin Chen
中科院分区:
数学2区
文献类型:
--
作者:
Gen Li;Yuting Wei;Yuejie Chi;Yuantao Gu;Yuxin Chen

文献摘要

被引文献

相似文献

Softmax策略梯度(PG)方法在softmax策略参数化下执行梯度上升,可以说是现代强化学习中策略优化的实际实现之一。对于折扣无限范围表格马尔可夫决策过程 (MDP),最近在建立 softmax PG 方法的全局收敛性以寻找接近最优策略方面取得了显着进展。然而,先前的结果未能明确描述收敛速度对显着参数的依赖性,例如状态空间的基数和有效范围,这两个参数都可能太大。在本文中,尽管假设可以进行精确的梯度计算,但我们对 softmax PG 方法的迭代复杂性传递了悲观的信息。具体来说,我们证明了具有步长的 softmax PG 方法可以收敛,即使存在良性策略初始化和适合探索的初始状态分布(以便分布失配系数不会太大)。这是通过在精心构建的仅包含三个操作的 MDP 上表征算法动态来实现的。我们的指数下限暗示有必要仔细调整更新规则或在加速 PG 方法中强制实施适当的正则化。
The softmax policy gradient (PG) method, which performs gradient ascent under softmax policy parameterization, is arguably one of the de facto implementations of policy optimization in modern reinforcement learning. For-discounted infinite-horizon tabular Markov decision processes (MDPs), remarkable progress has recently been achieved towards establishing global convergence of softmax PG methods in finding a near-optimal policy. However, prior results fall short of delineating clear dependencies of convergence rates on salient parameters such as the cardinality of the state spaceand the effective horizon, both of which could be excessively large. In this paper, we deliver a pessimistic message regarding the iteration complexity of softmax PG methods, despite assuming access to exact gradient computation. Specifically, we demonstrate that the softmax PG method with stepsizecan take{to} converge, even in the presence of a benign policy initialization and an initial state distribution amenable to exploration (so that the distribution mismatch coefficient is not exceedingly large). This is accomplished by characterizing the algorithmic dynamics over a carefully-constructed MDP containing only three actions. Our exponential lower bound hints at the necessity of carefully adjusting update rules or enforcing proper regularization in accelerating PG methods.