Detecting Network Effects: Randomizing Over Randomized Experiments

Detecting Network Effects: Randomizing Over Randomized Experiments
复制标题

检测网络效应:对随机实验进行随机化

DOI:
10.1145/3097983.3098192
复制
发表时间:
2017
期刊:
Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
E. Airoldi
E. Airoldi
中科院分区:
--
文献类型:
--
作者:
Martin Saveski;Jean Pouget;Guillaume Saint;Weitao Duan;Souvik Ghosh;Ya Xu;E. Airoldi

文献摘要

被引文献

相似文献

随机试验,或A/B检验,是评估新产品特征因果效应的标准方法,即,治疗。这些测试的有效性依赖于“稳定单位治疗值假设”(SUTVA),这意味着治疗只影响被治疗用户的行为,而不影响他们的连接行为。违反SUTVA,常见于表现出网络效应的特征,导致治疗因果效应的估计不准确。在本文中,我们利用一种新的实验设计来测试SUTVA是否成立,而不对治疗效果如何在治疗组和对照组之间溢出做出任何假设。为了实现这一点,我们同时运行完全随机化和基于聚类的随机化实验,然后比较所得估计值的差异。我们提出了一个统计检验测量这种差异的意义,并提供理论界的第一类错误率。我们为在大规模实验平台上实施我们的方法提供了实用的指导方针。重要的是,所提出的方法可以应用于网络不一定被观察到的设置,但如果可用的话,可以用于分析。最后,我们将此设计部署到LinkedIn的实验平台上,并将其应用于两个在线实验,突出了现实环境中标准A/B测试方法中网络效应和偏见的存在。
Randomized experiments, or A/B tests, are the standard approach for evaluating the causal effects of new product features, i.e., treatments. The validity of these tests rests on the "stable unit treatment value assumption" (SUTVA), which implies that the treatment only affects the behavior of treated users, and does not affect the behavior of their connections. Violations of SUTVA, common in features that exhibit network effects, result in inaccurate estimates of the causal effect of treatment. In this paper, we leverage a new experimental design for testing whether SUTVA holds, without making any assumptions on how treatment effects may spill over between the treatment and the control group. To achieve this, we simultaneously run both a completely randomized and a cluster-based randomized experiment, and then we compare the difference of the resulting estimates. We present a statistical test for measuring the significance of this difference and offer theoretical bounds on the Type I error rate. We provide practical guidelines for implementing our methodology on large-scale experimentation platforms. Importantly, the proposed methodology can be applied to settings in which a network is not necessarily observed but, if available, can be used in the analysis. Finally, we deploy this design to LinkedIn's experimentation platform and apply it to two online experiments, highlighting the presence of network effects and bias in standard A/B testing approaches in a real-world setting.