Model-Based Reinforcement Learning for Infinite-Horizon Discounted Constrained Markov Decision Processes

Model-Based Reinforcement Learning for Infinite-Horizon Discounted Constrained Markov Decision Processes
复制标题

DOI:
10.24963/ijcai.2021/347
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Aria HasanzadeZonuzy;D. Kalathil;S. Shakkottai
Aria HasanzadeZonuzy;D. Kalathil;S. Shakkottai
中科院分区:
其他
文献类型:
--
作者:
Aria HasanzadeZonuzy;D. Kalathil;S. Shakkottai

文献摘要

被引文献

相似文献

在许多现实世界的强化学习(RL)问题中,除了最大化目标之外,学习代理还必须保持一些必要的安全约束。我们制定的问题,学习一个安全的政策作为一个无限地平线折扣约束马尔可夫决策过程(CMDP)与未知的转移概率矩阵,其中的安全要求建模为约束预期的累积成本。我们提出了两种基于模型的约束强化学习(CRL)算法来学习安全策略,即(i)GM-CRL算法,其中算法可以访问生成模型,以及(ii)UC-CRL算法,其中算法使用上置信度风格在线探索方法学习模型。我们描述了这些算法的样本复杂度,即,在目标最大化和约束满足方面,确保高概率的期望精度水平所需的样本数量。
In many real-world reinforcement learning (RL) problems, in addition to maximizing the objective, the learning agent has to maintain some necessary safety constraints. We formulate the problem of learning a safe policy as an infinite-horizon discounted Constrained Markov Decision Process (CMDP) with an unknown transition probability matrix, where the safety requirements are modeled as constraints on expected cumulative costs. We propose two model-based constrained reinforcement learning (CRL) algorithms for learning a safe policy, namely, (i) GM-CRL algorithm, where the algorithm has access to a generative model, and (ii) UC-CRL algorithm, where the algorithm learns the model using an upper confidence style online exploration method. We characterize the sample complexity of these algorithms, i.e., the the number of samples needed to ensure a desired level of accuracy with high probability, both with respect to objective maximization and constraint satisfaction.