Assured RL: Reinforcement Learning with Almost Sure Constraints

Assured RL: Reinforcement Learning with Almost Sure Constraints
复制标题

Assured RL:具有几乎确定约束的强化学习

DOI:
--
复制
发表时间:
2020
期刊:
arXiv.org
影响因子:
--
通讯作者:
Enrique Mallada
Enrique Mallada
中科院分区:
--
文献类型:
--
作者:
Agustin Castellano;J. Bazerque;Enrique Mallada

文献摘要

参考文献

被引文献

相似文献

我们考虑了对状态转移和动作三元组具有几乎必然的约束的马尔可夫决策过程的最优策略寻找问题。我们定义的价值和行动价值函数满足基于障碍的分解,该分解允许独立于奖励过程识别可行的策略。我们证明了,在给定一个策略{pi}的情况下,证明某些状态-动作对在{pi}下是否导致可行轨迹等价于求解一个辅助问题,该辅助问题的目的是寻找执行不可行转移的概率。利用这种解释,我们开发了一种基于Q-学习的障碍学习算法,该算法识别这种不安全的状态-动作对。我们的分析激发了增强学习(RL)框架的需要,除了奖励之外,还需要一个额外的信号,在这里称为损伤函数,它提供可行性信息,并使具有无模型约束的RL问题的解决成为可能。此外,我们的障碍学习算法涵盖了现有的RL算法,如Q-学习和SARSA,使它们能够解决几乎肯定有约束的问题。
We consider the problem of finding optimal policies for a Markov Decision Process with almost sure constraints on state transitions and action triplets. We define value and action-value functions that satisfy a barrier-based decomposition which allows for the identification of feasible policies independently of the reward process. We prove that, given a policy {pi}, certifying whether certain state-action pairs lead to feasible trajectories under {pi} is equivalent to solving an auxiliary problem aimed at finding the probability of performing an unfeasible transition. Using this interpretation,we develop a Barrier-learning algorithm, based on Q-Learning, that identifies such unsafe state-action pairs. Our analysis motivates the need to enhance the Reinforcement Learning (RL) framework with an additional signal, besides rewards, called here damage function that provides feasibility information and enables the solution of RL problems with model-free constraints. Moreover, our Barrier-learning algorithm wraps around existing RL algorithms, such as Q-Learning and SARSA, giving them the ability to solve almost-surely constrained problems.
DOI: --
发表时间: 2020-03
期刊: --
影响因子: --
作者:
Dongsheng Ding;Xiaohan Wei;Zhuoran Yang;Zhaoran Wang;M. Jovanovi'c
通讯作者: Dongsheng Ding;Xiaohan Wei;Zhuoran Yang;Zhaoran Wang;M. Jovanovi'c