Predicting Counterfactuals from Large Historical Data and Small Randomized Trials

Predicting Counterfactuals from Large Historical Data and Small Randomized Trials
复制标题

DOI:
10.1145/3041021.3054190
复制
发表时间:
2016-10
期刊:
Proceedings of the 26th International Conference on World Wide Web Companion
影响因子:
--
通讯作者:
Nir Rosenfeld;Y. Mansour;E. Yom-Tov
Nir Rosenfeld;Y. Mansour;E. Yom-Tov
中科院分区:
其他
文献类型:
--
作者:
Nir Rosenfeld;Y. Mansour;E. Yom-Tov

文献摘要

被引文献

相似文献

当考虑使用一种新的治疗方法时,无论是药物还是搜索引擎排名算法,一个典型的问题是,它的性能是否会超过目前的治疗方法?回答这一反事实问题的传统方法是通过进行一项对照的随机试验来评估新治疗与传统治疗的效果。虽然这种方法在理论上确保了一个无偏的估计者,但它也有几个缺点,包括难以找到有代表性的实验总体以及运行随机试验的成本。此外,这样的试验忽略了大量可用的控制条件数据,这些数据原则上可以用于预测个性化效果的更困难的任务。在这篇文章中,我们提出了一个判别框架,用于从控制条件的大型数据集和来自比较新旧治疗的小型(可能不具代表性的)随机试验的数据中预测新治疗的结果。我们的学习目标需要对治疗做出最小的假设,对不同条件的结果之间的关系进行建模。这使我们不仅可以估计平均效应,而且还可以为小随机样本之外的例子生成单独的预测。我们通过三个方面的实验证明了我们的方法的实用性:搜索引擎操作,糖尿病患者的治疗,以及房屋的市场价值评估。我们的结果表明,我们的方法可以减少目前正在进行的随机对照实验的数量和规模,从而为实践者节省了大量的时间、金钱和精力。
When a new treatment is considered for use, whether a pharmaceutical drug or a search engine ranking algorithm, a typical question that arises is, will its performance exceed that of the current treatment? The conventional way to answer this counterfactual question is to estimate the effect of the new treatment in comparison to that of the conventional treatment by running a controlled, randomized experiment. While this approach theoretically ensures an unbiased estimator, it suffers from several drawbacks, including the difficulty in finding representative experimental populations as well as the cost of running randomized trials. Moreover, such trials neglect the huge quantities of available control-condition data, which in principle can be utilized for the harder task of predicting individualized effects. In this paper, we propose a discriminative framework for predicting the outcomes of a new treatment from a large dataset of the control condition and data from a small (and possibly unrepresentative) randomized trial comparing new and old treatments. Our learning objective, which requires minimal assumptions on the treatments, models the relation between the outcomes of the different conditions. This allows us to not only estimate mean effects but also to generate individual predictions for examples outside the small randomized sample. We demonstrate the utility of our approach through experiments in three areas: search engine operation, treatments to diabetes patients, and market value estimation of houses. Our results demonstrate that our approach can reduce the number and size of the currently performed randomized controlled experiments, thus saving significant time, money and effort on the part of practitioners.