Policy Learning with Constraints in Model-free Reinforcement Learning: A Survey

Policy Learning with Constraints in Model-free Reinforcement Learning: A Survey
复制标题

DOI:
10.24963/ijcai.2021/614
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Yongshuai Liu;A. Halev;Xin Liu
Yongshuai Liu;A. Halev;Xin Liu
中科院分区:
其他
文献类型:
--
作者:
Yongshuai Liu;A. Halev;Xin Liu

文献摘要

被引文献

相似文献

强化学习(RL)算法在模拟领域取得了巨大的成功。然而,这些算法通常不能直接应用于物理系统,特别是在需要满足约束的情况下(例如,为了确保安全或限制资源消耗)。在标准强化学习中,智能体被激励去探索任何策略,唯一的目标是最大化回报;然而,在真实的世界中,确保满足过程中的某些约束也是必要和必不可少的。在这篇文章中,我们概述了现有的解决无模型强化学习中约束的方法。我们建模的问题,学习的约束作为一个约束马尔可夫决策过程,并考虑两种主要类型的约束:累积和瞬时。我们总结了现有的方法,并讨论了它们的优点和缺点。为了评估政策的约束下的性能,我们引入了一套标准的基准和指标。我们还总结了当前方法的局限性,并提出了未来研究的悬而未决的问题。
Reinforcement Learning (RL) algorithms have had tremendous success in simulated domains. These algorithms, however, often cannot be directly applied to physical systems, especially in cases where there are constraints to satisfy (e.g. to ensure safety or limit resource consumption). In standard RL, the agent is incentivized to explore any policy with the sole goal of maximizing reward; in the real world, however, ensuring satisfaction of certain constraints in the process is also necessary and essential. In this article, we overview existing approaches addressing constraints in model-free reinforcement learning. We model the problem of learning with constraints as a Constrained Markov Decision Process and consider two main types of constraints: cumulative and instantaneous. We summarize existing approaches and discuss their pros and cons. To evaluate policy performance under constraints, we introduce a set of standard benchmarks and metrics. We also summarize limitations of current methods and present open questions for future research.