FLAME: A Fast Large-scale Almost Matching Exactly Approach to Causal Inference

FLAME: A Fast Large-scale Almost Matching Exactly Approach to Causal Inference
复制标题

DOI:
--
复制
发表时间:
2017-07
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Sudeepa Roy;C. Rudin;A. Volfovsky;Tianyu Wang-
Sudeepa Roy;C. Rudin;A. Volfovsky;Tianyu Wang-
中科院分区:
其他
文献类型:
--
作者:
Sudeepa Roy;C. Rudin;A. Volfovsky;Tianyu Wang-

文献摘要

被引文献

相似文献

因果推理中的一个经典问题是匹配问题,其中需要基于协变量信息将治疗单元与控制单元匹配。在这项工作中,我们提出了一种方法,计算高质量的几乎精确匹配的高维分类数据集。这种方法被称为FLAME(快速大规模几乎完全匹配),使用保持训练数据集学习匹配的距离度量。为了有效地执行大型数据集的匹配,FLAME利用了数据库管理领域中的查询处理的自然技术,并提供了两种FLAME实现:第一种使用SQL查询,第二种使用位向量技术。该算法首先构建最高质量的匹配(所有协变量的精确匹配),然后依次消除变量,以便尽可能多地精确匹配变量,同时仍然保持可解释的高质量匹配以及治疗组和对照组之间的平衡。我们利用这些高质量的匹配来估计条件平均治疗效果(CATE)。我们的实验表明,FLAME可扩展到具有数百万个观察结果的大型数据集,而现有的最先进方法无法做到这一点,并且它比其他匹配方法实现了显着更好的性能。
A classical problem in causal inference is that of matching, where treatment units need to be matched to control units based on covariate information. In this work, we propose a method that computes high quality almost-exact matches for high-dimensional categorical datasets. This method, called FLAME (Fast Large-scale Almost Matching Exactly), learns a distance metric for matching using a hold-out training data set. In order to perform matching efficiently for large datasets, FLAME leverages techniques that are natural for query processing in the area of database management, and two implementations of FLAME are provided: the first uses SQL queries and the second uses bit-vector techniques. The algorithm starts by constructing matches of the highest quality (exact matches on all covariates), and successively eliminates variables in order to match exactly on as many variables as possible, while still maintaining interpretable high-quality matches and balance between treatment and control groups. We leverage these high quality matches to estimate conditional average treatment effects (CATEs). Our experiments show that FLAME scales to huge datasets with millions of observations where existing state-of-the-art methods fail, and that it achieves significantly better performance than other matching methods.