A Provably-Efficient Model-Free Algorithm for Infinite-Horizon Average-Reward Constrained Markov Decision Processes

A Provably-Efficient Model-Free Algorithm for Infinite-Horizon Average-Reward Constrained Markov Decision Processes
复制标题

DOI:
10.1609/aaai.v36i4.20302
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Honghao Wei;Xin Liu;Lei Ying
Honghao Wei;Xin Liu;Lei Ying
中科院分区:
其他
文献类型:
--
作者:
Honghao Wei;Xin Liu;Lei Ying

文献摘要

被引文献

相似文献

本文介绍了无限制的加固学习(RL)算法,用于无限 - 霍尼平均奖励的马尔可夫决策过程(CMDPS)。考虑到足够大的学习范围K,提出的算法实现了sublerear的遗憾和零约束侵犯。界限取决于状态s的数量,动作A的数量和两个独立于学习视野K的常数。
This paper presents a model-free reinforcement learning (RL) algorithm for infinite-horizon average-reward Constrained Markov Decision Processes (CMDPs). Considering a learning horizon K, which is sufficiently large, the proposed algorithm achieves sublinear regret and zero constraint violation. The bounds depend on the number of states S, the number of actions A, and two constants which are independent of the learning horizon K.