Representation Balancing MDPs for Off-Policy Policy Evaluation

Representation Balancing MDPs for Off-Policy Policy Evaluation
复制标题

DOI:
--
复制
发表时间:
2018-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yao Liu;Omer Gottesman;Aniruddh Raghu;M. Komorowski;A. Faisal;F. Doshi-Velez;E. Brunskill
Yao Liu;Omer Gottesman;Aniruddh Raghu;M. Komorowski;A. Faisal;F. Doshi-Velez;E. Brunskill
中科院分区:
其他
文献类型:
--
作者:
Yao Liu;Omer Gottesman;Aniruddh Raghu;M. Komorowski;A. Faisal;F. Doshi-Velez;E. Brunskill

文献摘要

相似文献

研究了RL中的非政策策略评价问题。与之前的工作相比,我们考虑了如何准确地估计单个策略值和平均策略值。我们从最近的因果推理工作中汲取灵感,提出了一个新的有限样本泛化误差界,用于MDP模型的值估计。以这个上界为目标,我们开发了一种具有平衡表示的MDP模型的学习算法,并表明我们的方法可以在常见的合成基准和HIV治疗模拟领域产生更低的MSE。
We study the problem of off-policy policy evaluation (OPPE) in RL. In contrast to prior work, we consider how to estimate both the individual policy value and average policy value accurately. We draw inspiration from recent work in causal reasoning, and propose a new finite sample generalization error bound for value estimates from MDP models. Using this upper bound as an objective, we develop a learning algorithm of an MDP model with a balanced representation, and show that our approach can yield substantially lower MSE in common synthetic benchmarks and a HIV treatment simulation domain.