Post-Contextual-Bandit Inference

Post-Contextual-Bandit Inference
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
Advances in neural information processing systems
影响因子:
--
通讯作者:
Aurélien F. Bibaut;A. Chambaz;Maria Dimakopoulou;Nathan Kallus;M. Laan
Aurélien F. Bibaut;A. Chambaz;Maria Dimakopoulou;Nathan Kallus;M. Laan
中科院分区:
其他
文献类型:
--
作者:
Aurélien F. Bibaut;A. Chambaz;Maria Dimakopoulou;Nathan Kallus;M. Laan

文献摘要

被引文献

相似文献

上下文强盗算法正在越来越多地取代电子商务、医疗保健和政策制定中的非自适应A/B测试,因为它们既可以改善研究参与者的结果,又可以增加识别良好甚至最佳政策的机会。尽管如此,为了支持研究结束时对新干预措施的可信推断,我们仍然希望构建平均治疗效果、亚组效应或新政策价值的有效置信区间。然而,由上下文强盗算法收集的数据的自适应性质使得这变得困难:标准估计量不再是渐近正态分布的,并且经典置信区间无法提供正确的覆盖。虽然这已经解决了在非上下文环境中使用稳定的估计,上下文设置提出了独特的挑战,我们在本文中首次解决。我们提出了上下文自适应双重鲁棒(CADR)估计,第一个估计的政策价值是渐近正常的上下文自适应数据收集。在构建CADR的主要技术挑战是设计自适应和一致的条件标准差估计的稳定。使用57个OpenML数据集进行的大量数值实验表明,基于CADR的置信区间唯一地提供了正确的覆盖率。
Contextual bandit algorithms are increasingly replacing non-adaptive A/B tests in e-commerce, healthcare, and policymaking because they can both improve outcomes for study participants and increase the chance of identifying good or even best policies. To support credible inference on novel interventions at the end of the study, nonetheless, we still want to construct valid confidence intervals on average treatment effects, subgroup effects, or value of new policies. The adaptive nature of the data collected by contextual bandit algorithms, however, makes this difficult: standard estimators are no longer asymptotically normally distributed and classic confidence intervals fail to provide correct coverage. While this has been addressed in non-contextual settings by using stabilized estimators, the contextual setting poses unique challenges that we tackle for the first time in this paper. We propose the Contextual Adaptive Doubly Robust (CADR) estimator, the first estimator for policy value that is asymptotically normal under contextual adaptive data collection. The main technical challenge in constructing CADR is designing adaptive and consistent conditional standard deviation estimators for stabilization. Extensive numerical experiments using 57 OpenML datasets demonstrate that confidence intervals based on CADR uniquely provide correct coverage.