Causal inference and the data-fusion problem

Causal inference and the data-fusion problem
复制标题

DOI:
10.1073/pnas.1510507113
复制
发表时间:
2016-07-05
影响因子:
11.1
通讯作者:
Pearl, Judea
Pearl, Judea
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Bareinboim, Elias;Pearl, Judea

文献摘要

被引文献

相似文献

我们回顾了统一当前因果分析方法的概念、原则和工具,并关注大数据带来的新挑战。特别是,我们解决了数据融合的问题-将在异构条件下收集的多个数据集拼凑在一起(即,不同的群体、制度和抽样方法),以获得对感兴趣的查询的有效答案。多个异构数据集的可用性为大数据分析师提供了新的机会,因为可以从组合数据中获取的知识不可能单独来自任何单个来源。然而,在异构环境中出现的偏见需要新的分析工具。其中一些偏差,包括混杂,抽样选择和跨人群偏差,已被孤立地解决,主要是在限制参数模型。在这里,我们提出了一个一般的,非参数的框架来处理这些偏见,并最终在因果推理任务中的数据融合问题的理论解决方案。
We review concepts, principles, and tools that unify current approaches to causal analysis and attend to new challenges presented by big data. In particular, we address the problem of data fusion-piecing together multiple datasets collected under heterogeneous conditions (i.e., different populations, regimes, and sampling methods) to obtain valid answers to queries of interest. The availability of multiple heterogeneous datasets presents new opportunities to big data analysts, because the knowledge that can be acquired from combined data would not be possible from any individual source alone. However, the biases that emerge in heterogeneous environments require new analytical tools. Some of these biases, including confounding, sampling selection, and cross-population biases, have been addressed in isolation, largely in restricted parametric models. We here present a general, nonparametric framework for handling these biases and, ultimately, a theoretical solution to the problem of data fusion in causal inference tasks.