Off-Policy Evaluation for Large Action Spaces via Conjunct Effect Modeling

Off-Policy Evaluation for Large Action Spaces via Conjunct Effect Modeling
复制标题

DOI:
10.48550/arxiv.2305.08062
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuta Saito;Qingyang Ren;T. Joachims
Yuta Saito;Qingyang Ren;T. Joachims
中科院分区:
其他
文献类型:
--
作者:
Yuta Saito;Qingyang Ren;T. Joachims

文献摘要

被引文献

相似文献

我们研究了大型离散行动空间的上下文强盗策略的离策略评估(OPE),其中传统的重要性加权方法存在过度的方差。为了解决这个方差问题,我们提出了一种新的估计器,称为 OffCEM,它基于联合效应模型(CEM),一种将因果效应分解为集群效应和残差效应的新方法。 OffCEM 仅对行动集群应用重要性加权,并通过基于模型的奖励估计来解决残留因果效应。我们表明,所提出的估计量在称为局部正确性的新条件下是无偏的,它只要求残差效应模型保留每个集群内动作的相对预期奖励差异。为了最好地利用 CEM 和局部正确性,我们还提出了一种新的两步程序来执行基于模型的估计,最大限度地减少第一步中的偏差和第二步中的方差。我们发现,与一系列传统估计器相比,所得的 OffCEM 估计器大大改善了偏差和方差。实验表明,OffCEM 显着改进了 OPE,尤其是在存在许多操作的情况下。
We study off-policy evaluation (OPE) of contextual bandit policies for large discrete action spaces where conventional importance-weighting approaches suffer from excessive variance. To circumvent this variance issue, we propose a new estimator, called OffCEM, that is based on the conjunct effect model (CEM), a novel decomposition of the causal effect into a cluster effect and a residual effect. OffCEM applies importance weighting only to action clusters and addresses the residual causal effect through model-based reward estimation. We show that the proposed estimator is unbiased under a new condition, called local correctness, which only requires that the residual-effect model preserves the relative expected reward differences of the actions within each cluster. To best leverage the CEM and local correctness, we also propose a new two-step procedure for performing model-based estimation that minimizes bias in the first step and variance in the second step. We find that the resulting OffCEM estimator substantially improves bias and variance compared to a range of conventional estimators. Experiments demonstrate that OffCEM provides substantial improvements in OPE especially in the presence of many actions.