A Provably-Efficient Model-Free Algorithm for Infinite-Horizon Average-Reward Constrained Markov Decision Processes
A Provably-Efficient Model-Free Algorithm for Infinite-Horizon Average-Reward Constrained Markov Decision Processes
复制标题
DOI:
10.1609/aaai.v36i4.20302
复制
发表时间:
2022-06
期刊:
影响因子:
--
通讯作者:
Honghao Wei;Xin Liu;Lei Ying
中科院分区:
文献类型:
--
作者:
Honghao Wei;Xin Liu;Lei Ying
This paper presents a model-free reinforcement learning (RL) algorithm for infinite-horizon average-reward Constrained Markov Decision Processes (CMDPs). Considering a learning horizon K, which is sufficiently large, the proposed algorithm achieves sublinear regret and zero constraint violation. The bounds depend on the number of states S, the number of actions A, and two constants which are independent of the learning horizon K.