Towards A Complete Identification Algorithm for Missing Data Problems

Towards A Complete Identification Algorithm for Missing Data Problems
复制标题

面向缺失数据问题的完整识别算法

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
J. Robins
J. Robins
中科院分区:
--
文献类型:
--
作者:
I. Shpitser;J. Robins

文献摘要

被引文献

相似文献

因果推理的基本问题是缺少数据的问题——对两个假设的处理分配的反应的比较变得困难,因为对于每个实验单元,实际上只观察到一个处理分配。使用稳定的单位处理值假设和可忽略性,简单的识别结果导致将观测量和反事实量联系起来的因果推理已扩展到使用图形因果模型的完整一般理论[9,8,7],并且开发了观测数据的结果泛函的丰富估计理论。我们考虑相反观点的含义:缺失数据问题是因果推理的一种形式。我们认为从观测到的数据规律中识别完整数据规律的经典缺失数据问题是一个从事实变量的联合规律推断反事实变量的联合规律的问题。我们以一种与因果推理中的类似建模方法密切相关的方法,对图形模型中事实变量和反事实变量之间的关系进行编码,回顾了在该框架中开发的最新识别结果,并开发了一种新的算法,用于识别缺失数据和隐藏变量设置中的完整数据规律。我们的算法可以看作是ID算法的一个版本,用于识别适应缺失数据集特性的因果效应。我们的算法的完整性目前是一个开放的问题。
The fundamental problem of causal inference is a missing data problem – the comparison of responses to two hypothetical treatment assignments is made difficult because for every experimental unit, only one treatment assignment is actually observed. Simple identification results in causal inference that link observed and counterfactual quantities using the stable unit treatment value assumption and ignorability has been extended to a complete general theory using graphical causal models [9, 8, 7], and a rich estimation theory for resulting functionals of observed data has been developed. We consider the implications of the converse view: that missing data problems are a form of causal inference. We consider the classical missing data problem of identifying the full data law from the observed data law as a problem of inferring a joint law over counterfactual variables from a joint law over factual variables. We encode the relationship between the factual and counterfactual variables in graphical models, in an approach closely related to similar modeling approaches in causal inference, review recent identification results developed in this framework, and develop a new algorithm for identifying the full data law in settings with both missing data and hidden variables. Our algorithm can be viewed as a version of the ID algorithm for identifying causal effects adapted to peculiarities of the missing data setting. Completeness of our algorithm is currently an open problem.