Conservative and Adaptive Penalty for Model-Based Safe Reinforcement Learning

Conservative and Adaptive Penalty for Model-Based Safe Reinforcement Learning
复制标题

DOI:
10.1609/aaai.v36i5.20478
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Yecheng Jason Ma;Andrew Shen;O. Bastani;Dinesh Jayaraman
Yecheng Jason Ma;Andrew Shen;O. Bastani;Dinesh Jayaraman
中科院分区:
其他
文献类型:
--
作者:
Yecheng Jason Ma;Andrew Shen;O. Bastani;Dinesh Jayaraman

文献摘要

相似文献

现实世界中的强化学习(RL)代理人除了最大化奖励目标外,还必须满足安全限制。基于模型的RL算法有望减少不安全的现实世界动作:它们可以使用来自学习模型的模拟样本遵守所有约束的策略。但是,不完美的模型也可能导致现实世界中的违规行为,即使预计可以满足所有约束的行动。我们提出了保守和自适应惩罚(CAP),这是一个基于模型的安全RL框架,通过捕获模型不确定性并自适应利用它来平衡奖励和成本目标,以解决潜在的建模错误。首先,CAP使用基于不确定性的惩罚来膨胀预测的成本。从理论上讲,我们表明满足这种保守成本限制的政策在真实的环境中也是可行的。我们进一步表明,这可以保证RL培训期间所有中间解决方案的安全性。此外,CAP使用来自环境的真实成本反馈在培训期间适应这种罚款。我们评估了这种基于模型和基于图像的环境的基于模型的安全RL的保守和适应性惩罚方法。我们的结果表明,与先前的安全RL算法相比,样本效率的大幅增长,而违规行为少。代码可在以下网址找到:https://github.com/redrew/cap
Reinforcement Learning (RL) agents in the real world must satisfy safety constraints in addition to maximizing a reward objective. Model-based RL algorithms hold promise for reducing unsafe real-world actions: they may synthesize policies that obey all constraints using simulated samples from a learned model. However, imperfect models can result in real-world constraint violations even for actions that are predicted to satisfy all constraints. We propose Conservative and Adaptive Penalty (CAP), a model-based safe RL framework that accounts for potential modeling errors by capturing model uncertainty and adaptively exploiting it to balance the reward and the cost objectives. First, CAP inflates predicted costs using an uncertainty-based penalty. Theoretically, we show that policies that satisfy this conservative cost constraint are guaranteed to also be feasible in the true environment. We further show that this guarantees the safety of all intermediate solutions during RL training. Further, CAP adaptively tunes this penalty during training using true cost feedback from the environment. We evaluate this conservative and adaptive penalty-based approach for model-based safe RL extensively on state and image-based environments. Our results demonstrate substantial gains in sample-efficiency while incurring fewer violations than prior safe RL algorithms. Code is available at: https://github.com/Redrew/CAP