Offline Policy Optimization with Eligible Actions

Offline Policy Optimization with Eligible Actions
复制标题

DOI:
10.48550/arxiv.2207.00632
复制
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Yao Liu;Yannis Flet-Berliac;E. Brunskill
Yao Liu;Yannis Flet-Berliac;E. Brunskill
中科院分区:
其他
文献类型:
--
作者:
Yao Liu;Yannis Flet-Berliac;E. Brunskill

文献摘要

相似文献

由于在线学习在许多应用中可能是不可行的,因此fl的网络策略优化可能对许多现实世界的决策问题有很大的影响。重要性抽样及其变种是fl因策略评估中常用的估计器类型,这种估计器通常不需要对值函数或决策过程模型函数类的属性和表示能力进行假设。在本文中,我们识别了在优化重要性加权回报时的一个重要的超越fi现象,其中学习的策略可能本质上避免对部分初始状态空间进行对齐决策。我们通过一种新的按状态-邻域归一化约束,提出了一种在fi算法上避免这种情况的算法,并给出了该算法的一个理论上的fi证明。我们还指出了以前对这种方法的尝试的局限性。我们在一个受医疗保健启发的模拟器中测试了我们的算法,该模拟器是从真实医院和连续控制任务中收集的记录数据集。实验结果表明,与现有的批处理强化学习算法相比,本文提出的方法具有更好的测试性能和更少的fi训练时间。
Offline policy optimization could have a large im-pact on many real-world decision-making problems, as online learning may be infeasible in many applications. Importance sampling and its variants are a common used type of estimator in offline policy evaluation, and such estimators typically do not require assumptions on the properties and representational capabilities of value function or decision process model function classes. In this paper, we identify an important overfitting phe-nomenon in optimizing the importance weighted return, in which it may be possible for the learned policy to essentially avoid making aligned deci-sions for part of the initial state space. We propose an algorithm to avoid this overfitting through a new per-state-neighborhood normalization constraint, and provide a theoretical justification of the proposed algorithm. We also show the limita-tions of previous attempts to this approach. We test our algorithm in a healthcare-inspired simulator, a logged dataset collected from real hospitals and continuous control tasks. These experiments show the proposed method yields less overfitting and bet-ter test performance compared to state-of-the-art batch reinforcement learning algorithms.