Missing Data as a Causal and Probabilistic Problem

Missing Data as a Causal and Probabilistic Problem
复制标题

缺失数据作为因果和概率问题

DOI:
--
复制
发表时间:
2015
期刊:
Conference on Uncertainty in Artificial Intelligence
影响因子:
--
通讯作者:
J. Pearl
J. Pearl
中科院分区:
--
文献类型:
--
作者:
I. Shpitser;Karthika Mohan;J. Pearl

文献摘要

被引文献

相似文献

因果推断通常被表述为缺失数据问题-对于每个单元,只有对观察到的治疗分配的反应是已知的,对其他治疗分配的反应是未知的。在本文中,我们将[7]中表示缺失数据问题的匡威方法扩展到只允许对缺失指标进行干预的因果模型。我们进一步使用这种表示,以利用技术开发的因果效应识别的问题,给出了一个通用的标准的情况下,包含缺失变量的联合分布可以从实际观察到的数据中恢复,缺失机制的假设。这个标准是显着更普遍的比常用的“随机失踪”(MAR)的标准,并概括了过去的工作,也利用了图形表示的失踪。事实上,我们的标准与MAR的关系与识别因果效应的ID算法[22,18]和条件可验证性[13]之间的关系没有什么不同。
Causal inference is often phrased as a missing data problem - for every unit, only the response to observed treatment assignment is known, the response to other treatment assignments is not. In this paper, we extend the converse approach of [7] of representing missing data problems to causal models where only interventions on missingness indicators are allowed. We further use this representation to leverage techniques developed for the problem of identification of causal effects to give a general criterion for cases where a joint distribution containing missing variables can be recovered from data actually observed, given assumptions on missingness mechanisms. This criterion is significantly more general than the commonly used "missing at random" (MAR) criterion, and generalizes past work which also exploits a graphical representation of missingness. In fact, the relationship of our criterion to MAR is not unlike the relationship between the ID algorithm for identification of causal effects [22, 18], and conditional ignorability [13].