Synthetic Combinations: A Causal Inference Framework for Combinatorial Interventions

Synthetic Combinations: A Causal Inference Framework for Combinatorial Interventions
复制标题

综合组合:组合干预的因果推理框架

DOI:
10.48550/arxiv.2303.14226
复制
发表时间:
2023
期刊:
ArXiv
影响因子:
--
通讯作者:
S. Vijaykumar
S. Vijaykumar
中科院分区:
--
文献类型:
--
作者:
Abhineet Agarwal;Anish Agarwal;S. Vijaykumar

文献摘要

参考文献

被引文献

相似文献

考虑一个有 $N$ 异质单位和 $p$ 干预的环境。我们的目标是了解这些 $p$ 干预措施的任意组合的特定于单元的潜在结果,即 $N \times 2^p$ 因果参数。选择干预措施组合是各种应用中自然出现的问题,例如因子设计实验、推荐引擎、医学组合疗法、联合分析等。随着 $N$ 和 $p$ 的增长,运行 $N \times 2^p$ 实验来估计各种参数可能很昂贵和/或不可行。此外,观察数据可能存在混淆,即在组合下是否看到一个单元与其在该组合下的潜在结果相关。为了应对这些挑战,我们提出了一种新颖的潜在因素模型,该模型强加了跨单元的结构(即潜在结果的矩阵大约为$r$)和干预组合(即潜在结果的傅里叶展开中的系数大约为$s$稀疏)。尽管存在未观察到的混杂因素,我们仍为所有 $N \times 2^p$ 参数建立了标识。我们提出了一种估计程序,即综合组合,并确定它在观测模式的精确条件下是有限样本一致的和渐近正态的。我们的结果意味着给定 $\text{poly}(r) \times \left( N + s^2p\right)$ 观测值的估计是一致的,而以前的方法将样本复杂度缩放为 $\min(N \times s^2p, \ \ \text{poly(r)} \times (N + 2^p))$。我们使用合成组合来提出数据高效的实验设计。根据经验,综合组合在电影推荐的真实数据集上优于竞争方法。最后,我们将分析扩展到进行因果推断,其中干预是对 $p$ 项目(例如排名)的排列。
Consider a setting where there are $N$ heterogeneous units and $p$ interventions. Our goal is to learn unit-specific potential outcomes for any combination of these $p$ interventions, i.e., $N \times 2^p$ causal parameters. Choosing a combination of interventions is a problem that naturally arises in a variety of applications such as factorial design experiments, recommendation engines, combination therapies in medicine, conjoint analysis, etc. Running $N \times 2^p$ experiments to estimate the various parameters is likely expensive and/or infeasible as $N$ and $p$ grow. Further, with observational data there is likely confounding, i.e., whether or not a unit is seen under a combination is correlated with its potential outcome under that combination. To address these challenges, we propose a novel latent factor model that imposes structure across units (i.e., the matrix of potential outcomes is approximately rank $r$), and combinations of interventions (i.e., the coefficients in the Fourier expansion of the potential outcomes is approximately $s$ sparse). We establish identification for all $N \times 2^p$ parameters despite unobserved confounding. We propose an estimation procedure, Synthetic Combinations, and establish it is finite-sample consistent and asymptotically normal under precise conditions on the observation pattern. Our results imply consistent estimation given $\text{poly}(r) \times \left( N + s^2p\right)$ observations, while previous methods have sample complexity scaling as $\min(N \times s^2p, \ \ \text{poly(r)} \times (N + 2^p))$. We use Synthetic Combinations to propose a data-efficient experimental design. Empirically, Synthetic Combinations outperforms competing approaches on a real-world dataset on movie recommendations. Lastly, we extend our analysis to do causal inference where the intervention is a permutation over $p$ items (e.g., rankings).
从潜在结果的角度进行基于回归的因果分析。
DOI: 10.1515/jem-2018-0030
发表时间: 2020
影响因子: --
作者:
Terza,JosephV
通讯作者: Terza,JosephV
使用决策树进行大规模预测
DOI: 10.1080/01621459.2022.2126782
发表时间: 2023
影响因子: 3.7
作者:
Klusowski, Jason M.;Tian, Peter M.
通讯作者: Tian, Peter M.