Learning Safe Policies via Primal-Dual Methods

Learning Safe Policies via Primal-Dual Methods
复制标题

通过原始对偶方法学习安全策略

DOI:
--
复制
发表时间:
2019
期刊:
IEEE Conference on Decision and Control
影响因子:
--
通讯作者:
Alejandro Ribeiro
Alejandro Ribeiro
中科院分区:
--
文献类型:
--
作者:
Santiago Paternain;Miguel Calvo;Luiz F. O. Chamon;Alejandro Ribeiro

文献摘要

被引文献

相似文献

在本文中,我们研究了安全策略的学习在强化学习问题的设置。也就是说,我们的目标是控制马尔可夫决策过程(MDP),我们不知道其转移概率,但我们可以通过实验获得样本轨迹。我们将安全性定义为代理在每个时间实例中都以高概率保持在所需的安全集合中。因此,我们考虑一个受约束的MDP的约束是概率。由于在强化学习框架中很难解决这些约束,我们提出了一个遍历放松的问题。尽管如此,这种放宽使我们能够对由此产生的政策提供安全保证。为了计算这些政策,我们资源的随机原始对偶方法。我们测试所提出的方法在网格世界中的导航任务。数值结果表明,我们的算法是能够动态地适应环境和所需的安全级别的政策。
In this paper, we study the learning of safe policies in the setting of reinforcement learning problems. This is, we aim to control a Markov Decision Process (MDP) of which we do not know the transition probabilities, but we have access to sample trajectories through experiments. We define safety as the agent remaining in a desired safe set with high probability for every time instance. We therefore consider a constrained MDP where the constraints are probabilistic. Due to the difficulty of addressing these constraints in a reinforcement learning framework, we propose an ergodic relaxation of the problem. Nonetheless, this relaxation is such that we are able to provide safety guarantees on the resulting policies. To compute these policies, we resource to a stochastic primal-dual method. We test the proposed approach in a navigation task in a grid world. The numerical results show that our algorithm is capable of dynamically adapting the policy to the environment and the required safety levels.