Boosted Decision Tree Regression Adjustment for Variance Reduction in Online Controlled Experiments

Boosted Decision Tree Regression Adjustment for Variance Reduction in Online Controlled Experiments
复制标题

在线控制实验中减少方差的增强决策树回归调整

DOI:
--
复制
发表时间:
2016
期刊:
Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
P. Serdyukov
P. Serdyukov
中科院分区:
--
文献类型:
--
作者:
Alexey Poyarkov;Alexey Drutsa;A. Khalyavin;Gleb Gusev;P. Serdyukov

文献摘要

被引文献

相似文献

如今,大多数领先的Web服务的开发都是由在线实验控制的,这些实验对它们的稳定更新流进行定性和量化,每天实现一千多个并发实验。尽管越来越需要运行更多的实验,但这些服务的用户流量有限。这种情况导致的问题,找到一个新的或改进现有的关键性能指标具有较高的灵敏度和较低的方差。我们专注于减少方差的问题,广泛用于A/B测试的Web服务的用户忠诚度的参与度指标。我们开发了一个通用的框架,是基于评估的平均值之间的差异的实际和近似值的关键性能指标(而不是这个指标的平均值)。一方面,它使我们能够将最先进的技术广泛用于临床和社会研究的随机实验,但在在线评估中的使用有限。另一方面,我们提出了一类新的方法,基于先进的机器学习算法,包括集成的决策树,据我们所知,还没有被应用到更早的方差减少的问题。我们在Yandex上运行的一组非常大的真实的大规模A/B实验上验证了方差减少方法,用于不同的用户忠诚度参与指标。我们最好的方法证明了$63\%$平均方差减少(这相当于63%节省用户流量),并检测治疗效果在$2$倍以上的A/B实验。
Nowadays, the development of most leading web services is controlled by online experiments that qualify and quantify the steady stream of their updates achieving more than a thousand concurrent experiments per day. Despite the increasing need for running more experiments, these services are limited in their user traffic. This situation leads to the problem of finding a new or improving existing key performance metric with a higher sensitivity and lower variance. We focus on the problem of variance reduction for engagement metrics of user loyalty that are widely used in A/B testing of web services. We develop a general framework that is based on evaluation of the mean difference between the actual and the approximated values of the key performance metric (instead of the mean of this metric). On the one hand, it allows us to incorporate the state-of-the-art techniques widely used in randomized experiments of clinical and social research, but limitedly used in online evaluation. On the other hand, we propose a new class of methods based on advanced machine learning algorithms, including ensembles of decision trees, that, to the best of our knowledge, have not been applied earlier to the problem of variance reduction. We validate the variance reduction approaches on a very large set of real large-scale A/B experiments run at Yandex for different engagement metrics of user loyalty. Our best approach demonstrates $63\%$ average variance reduction (which is equivalent to 63% saved user traffic) and detects the treatment effect in $2$ times more A/B experiments.