Variance-Minimizing Augmentation Logging for Counterfactual Evaluation in Contextual Bandits

Variance-Minimizing Augmentation Logging for Counterfactual Evaluation in Contextual Bandits
复制标题

DOI:
10.1145/3539597.3570452
复制
发表时间:
2023-02
期刊:
Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining
影响因子:
--
通讯作者:
Aaron David Tucker;T. Joachims
Aaron David Tucker;T. Joachims
中科院分区:
其他
文献类型:
--
作者:
Aaron David Tucker;T. Joachims

文献摘要

相似文献

离线A/B测试和反事实学习的方法正在搜索和推荐系统中迅速被采用,因为它们允许有效地重新使用现有的日志数据。然而,单独使用现有日志数据存在基本限制,因为当日志记录策略与正在评估的目标策略非常不同时,这些方法中常用的反事实估计器可能会有很大的偏差和很大的方差。为了克服这一限制,我们探索了如何设计数据收集策略的问题,以最有效地利用学习和评估的额外观察来增强现有的强盗反馈数据集。为此,本文引入了最小方差增强日志记录(MVAL),这是一种构造日志记录策略的方法,该方法可以最小化下游评估或学习问题的方差。我们探索了多种方法来有效地计算MVAL策略,并发现它们在降低估计量的方差方面比幼稚方法要有效得多。
Methods for offline A/B testing and counterfactual learning are seeing rapid adoption in search and recommender systems, since they allow efficient reuse of existing log data. However, there are fundamental limits to using existing log data alone, since the counterfactual estimators that are commonly used in these methods can have large bias and large variance when the logging policy is very different from the target policy being evaluated. To overcome this limitation, we explore the question of how to design data-gathering policies that most effectively augment an existing dataset of bandit feedback with additional observations for both learning and evaluation. To this effect, this paper introduces Minimum Variance Augmentation Logging (MVAL), a method for constructing logging policies that minimize the variance of the downstream evaluation or learning problem. We explore multiple approaches to computing MVAL policies efficiently, and find that they can be substantially more effective in decreasing the variance of an estimator than naïve approaches.