IPO: Interior-point Policy Optimization under Constraints

IPO: Interior-point Policy Optimization under Constraints
复制标题

DOI:
10.1609/aaai.v34i04.5932
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yongshuai Liu;J. Ding;Xin Liu
Yongshuai Liu;J. Ding;Xin Liu
中科院分区:
其他
文献类型:
--
作者:
Yongshuai Liu;J. Ding;Xin Liu

文献摘要

相似文献

在本文中,我们研究了强化学习(RL)算法来解决现实世界的决策问题的目标是最大化的长期回报,以及满足累积约束。我们提出了一种新的一阶策略优化方法,内点策略优化(IPO),它增加了对数障碍函数的目标,启发了邻点方法。我们提出的方法是很容易实现的性能保证,并可以处理一般类型的累积多约束设置。我们进行广泛的评估,将我们的方法与最先进的基线进行比较。我们的算法优于基线算法,在奖励最大化和约束满足。
In this paper, we study reinforcement learning (RL) algorithms to solve real-world decision problems with the objective of maximizing the long-term reward as well as satisfying cumulative constraints. We propose a novel first-order policy optimization method, Interior-point Policy Optimization (IPO), which augments the objective with logarithmic barrier functions, inspired by the interior-point method. Our proposed method is easy to implement with performance guarantees and can handle general types of cumulative multi-constraint settings. We conduct extensive evaluations to compare our approach with state-of-the-art baselines. Our algorithm outperforms the baseline algorithms, in terms of reward maximization and constraint satisfaction.