IPO: Interior-point Policy Optimization under Constraints
IPO: Interior-point Policy Optimization under Constraints
复制标题
DOI:
10.1609/aaai.v34i04.5932
复制
发表时间:
2019-10
期刊:
影响因子:
--
通讯作者:
Yongshuai Liu;J. Ding;Xin Liu
中科院分区:
文献类型:
--
作者:
Yongshuai Liu;J. Ding;Xin Liu
In this paper, we study reinforcement learning (RL) algorithms to solve real-world decision problems with the objective of maximizing the long-term reward as well as satisfying cumulative constraints. We propose a novel first-order policy optimization method, Interior-point Policy Optimization (IPO), which augments the objective with logarithmic barrier functions, inspired by the interior-point method. Our proposed method is easy to implement with performance guarantees and can handle general types of cumulative multi-constraint settings. We conduct extensive evaluations to compare our approach with state-of-the-art baselines. Our algorithm outperforms the baseline algorithms, in terms of reward maximization and constraint satisfaction.