Policy gradient primal-dual mirror descent for constrained MDPs with large state spaces
Policy gradient primal-dual mirror descent for constrained MDPs with large state spaces
复制标题
DOI:
10.1109/cdc51059.2022.9992419
复制
发表时间:
2022-12
期刊:
影响因子:
--
通讯作者:
Dongsheng Ding;M. Jovanović
中科院分区:
文献类型:
--
作者:
Dongsheng Ding;M. Jovanović
We study constrained sequential decision-making problems modeled by constrained Markov decision processes with potentially infinite state spaces. We propose a Bregman distance-based direct policy search method – policy gradient primal-dual mirror descent – which includes the natural policy primal-dual method and the projected policy primal-dual method as two special cases. When the exact gradient is known, we prove dimension-free global convergence with a sublinear rate in both optimality gap and constraint violation. When the exact gradient is not available, we instantiate our algorithm in the linear function approximation setting and establish sample complexity guarantees. The introduction of the Bregman-distance regularizers enjoys the dimension-free property with applicability to large-scale spaces, the first of its kind in the constrained RL literature.