Balanced Off-Policy Evaluation in General Action Spaces

Balanced Off-Policy Evaluation in General Action Spaces
复制标题

一般行动空间中的平衡非政策评估

DOI:
--
复制
发表时间:
2019
期刊:
International Conference on Artificial Intelligence and Statistics
影响因子:
--
通讯作者:
Drew Dimmery
Drew Dimmery
中科院分区:
--
文献类型:
--
作者:
A. Sondhi;D. Arbour;Drew Dimmery

文献摘要

参考文献

被引文献

相似文献

用于上下文强盗的非策略评估的重要性采样权重的估计通常导致不平衡-加权后的状态-动作对的期望分布与实际分布之间的不匹配。在这项工作中,我们提出了平衡的政策评估(B-OPE),一个通用的方法来估计权重,最大限度地减少这种不平衡。这些权重的估计减少到一个二进制分类问题,无论动作类型。我们表明,最大限度地减少分类器的风险意味着最小化的不平衡所需的反事实分布的状态-动作对。分类器损失与非策略估计的误差相关,从而允许轻松调整超参数。我们提供的实验证据表明,B-OPE改善了基于权重的方法,离线政策评估离散和连续的行动空间。
Estimation of importance sampling weights for off-policy evaluation of contextual bandits often results in imbalance - a mismatch between the desired and the actual distribution of state-action pairs after weighting. In this work we present balanced off-policy evaluation (B-OPE), a generic method for estimating weights which minimize this imbalance. Estimation of these weights reduces to a binary classification problem regardless of action type. We show that minimizing the risk of the classifier implies minimization of imbalance to the desired counterfactual distribution of state-action pairs. The classifier loss is tied to the error of the off-policy estimate, allowing for easy tuning of hyperparameters. We provide experimental evidence that B-OPE improves weighting-based approaches for offline policy evaluation in both discrete and continuous action spaces.
DOI: --
发表时间: 2018
期刊: Proceedings of the 21st International Conference on Artificial Intelligence and Statistics (AISTATS
影响因子: --
作者:
Kallus, Nathan;Zhou, Angela
通讯作者: Zhou, Angela
平衡的政策评估和学习
DOI: --
发表时间: 2018
期刊: Advances in neural information processing systems
影响因子: --
作者:
Kallus, Nathan
通讯作者: Kallus, Nathan
DOI: 10.1214/07-sts227b
发表时间: 2007-01-01
期刊: Statistical science : a review journal of the Institute of Mathematical Statistics
影响因子: --
作者:
Tsiatis, Anastasios A;Davidian, Marie
通讯作者: Davidian, Marie