Policy gradient primal-dual mirror descent for constrained MDPs with large state spaces

Policy gradient primal-dual mirror descent for constrained MDPs with large state spaces
复制标题

DOI:
10.1109/cdc51059.2022.9992419
复制
发表时间:
2022-12
期刊:
2022 IEEE 61st Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Dongsheng Ding;M. Jovanović
Dongsheng Ding;M. Jovanović
中科院分区:
其他
文献类型:
--
作者:
Dongsheng Ding;M. Jovanović

文献摘要

相似文献

我们研究了约束马尔可夫决策过程与潜在无限状态空间建模的约束序列决策问题。我们提出了一种基于Bregman距离的直接策略搜索方法-策略梯度原对偶镜像下降法-其中包括自然策略原对偶方法和投影策略原对偶方法作为两种特殊情况。当精确梯度已知时,我们证明了无量纲的全局收敛性,在最优性间隙和约束违反下都具有次线性速度。当精确梯度不可用时,我们在线性函数近似设置中实例化我们的算法,并建立样本复杂度保证。Bregman距离正则化子的引入具有适用于大规模空间的无量纲性质,这在约束RL文献中是第一次。
We study constrained sequential decision-making problems modeled by constrained Markov decision processes with potentially infinite state spaces. We propose a Bregman distance-based direct policy search method – policy gradient primal-dual mirror descent – which includes the natural policy primal-dual method and the projected policy primal-dual method as two special cases. When the exact gradient is known, we prove dimension-free global convergence with a sublinear rate in both optimality gap and constraint violation. When the exact gradient is not available, we instantiate our algorithm in the linear function approximation setting and establish sample complexity guarantees. The introduction of the Bregman-distance regularizers enjoys the dimension-free property with applicability to large-scale spaces, the first of its kind in the constrained RL literature.