Information-Directed Policy Search in Sparse-Reward Settings via the Occupancy Information Ratio
Information-Directed Policy Search in Sparse-Reward Settings via the Occupancy Information Ratio
复制标题
DOI:
10.1109/ciss56502.2023.10089655
复制
发表时间:
2023-03
期刊:
影响因子:
--
通讯作者:
Wesley A. Suttle;Alec Koppel;Ji Liu
中科院分区:
文献类型:
--
作者:
Wesley A. Suttle;Alec Koppel;Ji Liu
This paper examines a new measure of the exploration/exploitation trade-off in reinforcement learning (RL) called the occupancy information ratio (OIR). To this end, the paper derives the Information-Directed Actor-Critic (IDAC) algorithm for solving the OIR problem, provides an overview of the rich theory underlying IDAC and related OIR policy gradient methods, and experimentally investigates the advantages of such methods. The central contribution of this paper is to provide empirical evidence that, due to the form of the OIR objective, IDAC enjoys superior performance over vanilla RL methods in sparse-reward environments.