Scalable Probabilistic Causal Structure Discovery

Scalable Probabilistic Causal Structure Discovery
复制标题

DOI:
10.24963/ijcai.2018/709
复制
发表时间:
2018-07
影响因子:
1
通讯作者:
Dhanya Sridhar;J. Pujara;L. Getoor
Dhanya Sridhar;J. Pujara;L. Getoor
中科院分区:
经济学4区
文献类型:
--
作者:
Dhanya Sridhar;J. Pujara;L. Getoor

文献摘要

相似文献

复杂的因果网络构成了许多现实世界问题的基础,从基因之间的调控相互作用到用于理解气候变化的环境模式。计算方法试图利用观测数据和领域知识来推断这些因果网络。在本文中,我们确定了推断科学发现因果网络结构的三个关键要求:(1)对观测测量中的噪声具有健壮性;(2)处理数百个变量的可伸缩性;以及(3)编码领域知识和其他结构约束的灵活性。我们首先将联合概率因果结构发现问题形式化。我们开发了一种使用概率软逻辑(PSL)的方法,该方法利用多个统计测试,支持对数百个变量的有效优化,并且可以很容易地纳入结构约束,包括不完善的领域知识。我们将我们的方法与在生物和合成数据集上进行了充分研究的多种方法进行了比较,结果显示,在现实环境中,F1分数比最佳执行基线提高了20%。
Complex causal networks underlie many real-world problems, from the regulatory interactions between genes to the environmental patterns used to understand climate change. Computational methods seek to infer these causal networks using observational data and domain knowledge. In this paper, we identify three key requirements for inferring the structure of causal networks for scientific discovery: (1) robustness to noise in observed measurements; (2) scalability to handle hundreds of variables; and (3) flexibility to encode domain knowledge and other structural constraints. We first formalize the problem of joint probabilistic causal structure discovery. We develop an approach using probabilistic soft logic (PSL) that exploits multiple statistical tests, supports efficient optimization over hundreds of variables, and can easily incorporate structural constraints, including imperfect domain knowledge. We compare our method against multiple well-studied approaches on biological and synthetic datasets, showing improvements of up to 20% in F1-score over the best performing baseline in realistic settings.