A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision Processes

A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision Processes
复制标题

DOI:
--
复制
发表时间:
2021-11
影响因子:
2.2
通讯作者:
C. Shi;Masatoshi Uehara;Jiawei Huang;Nan Jiang
C. Shi;Masatoshi Uehara;Jiawei Huang;Nan Jiang
中科院分区:
计算机科学4区
文献类型:
--
作者:
C. Shi;Masatoshi Uehara;Jiawei Huang;Nan Jiang

文献摘要

相似文献

我们考虑部分可观察马尔可夫决策过程(POMDP)中的离策略评估(OPE),其中评估策略仅取决于可观察变量,而行为策略取决于不可观察潜在变量。现有的工作要么假设没有不可测量的混杂因素,要么专注于观察空间和状态空间都是表格的设置。在这项工作中,我们首先通过引入连接目标策略值和观察到的数据分布的桥接函数,提出了具有潜在混杂因素的 POMDP 中 OPE 的新颖识别方法。接下来,我们提出用于学习这些桥函数的极小极大估计方法,并基于这些估计的桥函数构造三个估计器,对应于基于价值函数的估计器、边缘重要性采样估计器和双鲁棒估计器。我们的建议允许一般函数逼近,因此适用于具有连续或大观察/状态空间的设置。详细研究了所提出的估计量的非渐近和渐近性质。
We consider off-policy evaluation (OPE) in Partially Observable Markov Decision Processes (POMDPs), where the evaluation policy depends only on observable variables and the behavior policy depends on unobservable latent variables. Existing works either assume no unmeasured confounders, or focus on settings where both the observation and the state spaces are tabular. In this work, we first propose novel identification methods for OPE in POMDPs with latent confounders, by introducing bridge functions that link the target policy's value and the observed data distribution. We next propose minimax estimation methods for learning these bridge functions, and construct three estimators based on these estimated bridge functions, corresponding to a value function-based estimator, a marginalized importance sampling estimator, and a doubly-robust estimator. Our proposal permits general function approximation and is thus applicable to settings with continuous or large observation/state spaces. The nonasymptotic and asymptotic properties of the proposed estimators are investigated in detail.