Information-Directed Policy Search in Sparse-Reward Settings via the Occupancy Information Ratio

Information-Directed Policy Search in Sparse-Reward Settings via the Occupancy Information Ratio
复制标题

DOI:
10.1109/ciss56502.2023.10089655
复制
发表时间:
2023-03
期刊:
2023 57th Annual Conference on Information Sciences and Systems (CISS)
影响因子:
--
通讯作者:
Wesley A. Suttle;Alec Koppel;Ji Liu
Wesley A. Suttle;Alec Koppel;Ji Liu
中科院分区:
其他
文献类型:
--
作者:
Wesley A. Suttle;Alec Koppel;Ji Liu

文献摘要

相似文献

本文研究了强化学习(RL)中探索/开发权衡的一种新度量,称为占用信息比(OIR)。为此,本文推导了用于解决OIR问题的信息指导的演员-评论家(IDAC)算法,概述了IDAC算法和相关OIR策略梯度方法的丰富理论基础,并通过实验研究了这些方法的优势。本文的主要贡献是提供经验证据,由于OIR目标的形式,IDAC享有上级性能优于香草RL方法在稀疏奖励环境。
This paper examines a new measure of the exploration/exploitation trade-off in reinforcement learning (RL) called the occupancy information ratio (OIR). To this end, the paper derives the Information-Directed Actor-Critic (IDAC) algorithm for solving the OIR problem, provides an overview of the rich theory underlying IDAC and related OIR policy gradient methods, and experimentally investigates the advantages of such methods. The central contribution of this paper is to provide empirical evidence that, due to the form of the OIR objective, IDAC enjoys superior performance over vanilla RL methods in sparse-reward environments.