Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Primal-Dual Approach

Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Primal-Dual Approach
复制标题

通过原始对偶方法实现约束强化学习的零约束违规

DOI:
--
复制
发表时间:
2021
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
V. Aggarwal
V. Aggarwal
中科院分区:
--
文献类型:
--
作者:
Qinbo Bai;A. S. Bedi;Mridul Agarwal;Alec Koppel;V. Aggarwal

文献摘要

参考文献

被引文献

相似文献

强化学习广泛应用于需要在与环境交互时执行顺序决策的应用中。当决策要求包括满足一些安全约束时,问题变得更具挑战性。该问题在数学上表示为约束马尔可夫决策过程(CMDP)。在文献中,各种算法可用于以无模型的方式解决CMDP问题,以实现ε-最优的累积奖励,并具有非线性可行的策略。一个ε可行的策略意味着它遭受约束违反。这里的一个重要问题是,我们是否可以在零约束违反的情况下实现ε最优累积奖励。为了实现这一目标,我们提倡使用随机原始-对偶方法来解决CMDP问题,并提出了一种保守的随机原始-对偶算法(CSPDA),该算法具有O(1/ε ^2)的样本复杂度,可以在零约束违反的情况下实现ε-最优累积奖励。在之前的工作中,违反约束为零的ε最优策略的最佳可用样本复杂度为O(1/ε ^5)。因此,与现有技术相比,所提出的算法提供了显著的改进。
Reinforcement learning is widely used in applications where one needs to perform sequential decisions while interacting with the environment. The problem becomes more challenging when the decision requirement includes satisfying some safety constraints. The problem is mathematically formulated as constrained Markov decision process (CMDP). In the literature, various algorithms are available to solve CMDP problems in a model-free manner to achieve epsilon-optimal cumulative reward with epsilon feasible policies. An epsilon-feasible policy implies that it suffers from constraint violation. An important question here is whether we can achieve epsilon-optimal cumulative reward with zero constraint violations or not. To achieve that, we advocate the use of a randomized primal-dual approach to solve the CMDP problems and propose a conservative stochastic primal-dual algorithm (CSPDA) which is shown to exhibit O(1/epsilon^2) sample complexity to achieve epsilon-optimal cumulative reward with zero constraint violations. In the prior works, the best available sample complexity for the epsilon-optimal policy with zero constraint violation is O(1/epsilon^5). Hence, the proposed algorithm provides a significant improvement compared to the state of the art.
DOI: 10.1609/aaai.v35i9.16979
发表时间: 2020-09
期刊: ArXiv
影响因子: --
作者:
K. C. Kalagarla;Rahul Jain;P. Nuzzo
通讯作者: K. C. Kalagarla;Rahul Jain;P. Nuzzo
DOI: --
发表时间: 2020-03
期刊: --
影响因子: --
作者:
Dongsheng Ding;Xiaohan Wei;Zhuoran Yang;Zhaoran Wang;M. Jovanovi'c
通讯作者: Dongsheng Ding;Xiaohan Wei;Zhuoran Yang;Zhaoran Wang;M. Jovanovi'c
DOI: 10.1007/s00239-010-9414-3
发表时间: 2011-02-01
影响因子: 3.9
作者:
Riepsamen, Angelique H.;Gibson, Tracey;Dowton, Mark
通讯作者: Dowton, Mark
贴现 MDP 的近极小极大最优强化学习
DOI: --
发表时间: 2021
期刊: Advances in neural information processing systems
影响因子: --
作者:
He, Jiafan;Zhou, Dongruo;Gu, Quanquan
通讯作者: Gu, Quanquan