Improved baselines for causal structure learning on interventional data

Improved baselines for causal structure learning on interventional data
复制标题

DOI:
10.1007/s11222-023-10257-9
复制
发表时间:
2023-10-01
影响因子:
2.2
通讯作者:
Mukherjee,Sach
Mukherjee,Sach
中科院分区:
数学2区
文献类型:
--
作者:
Richter,Robin;Bhamidi,Shankar;Mukherjee,Sach

文献摘要

相似文献

因果结构学习(CSL)是指从数据中估计因果图。因果版本的工具,如ROC曲线在CSL方法的经验评估中发挥了重要作用,并且经常将性能与“随机”基线(如ROC分析中的对角线)进行比较。然而,这样的基线没有考虑到从图形上下文产生的约束,因此可能代表“低条”。在本文中,在系统生物学的例子的动机,我们专注于评估CSL方法的多变量数据的图形结构的一部分是已知的,通过干预实验。对于这种设置,我们提出了一类新的基线,称为基于图的预测器(GBP)。与“随机”基线相比,GBP利用已知的图结构,利用简单的图属性来提供改进的基线,以比较CSL方法。我们一般讨论GBP,并在传递闭图的背景下提供了详细的研究,为这种设置引入了两个概念上简单的基线,观察到的程度预测(OIP)和传递性假设预测(TAP)。虽然前者是简单的计算,对于后者,我们提出了几种模拟策略。此外,我们研究和比较所提出的预测理论,包括一个结果表明,OIP优于预期的“随机”基线上的一个子类的潜在网络模型具有正相关的边缘概率。使用模拟和真实的生物数据,我们表明,建议的GBP优于随机基线在实践中,往往大幅。一些GBP甚至优于标准CSL方法(同时在实践中计算便宜)。我们的研究结果提供了一种新的方法来评估CSL方法的介入数据。
Causal structure learning (CSL) refers to the estimation of causal graphs from data. Causal versions of tools such as ROC curves play a prominent role in empirical assessment of CSL methods and performance is often compared with “random” baselines (such as the diagonal in an ROC analysis). However, such baselines do not take account of constraints arising from the graph context and hence may represent a “low bar”. In this paper, motivated by examples in systems biology, we focus on assessment of CSL methods for multivariate data where part of the graph structure is known via interventional experiments. For this setting, we put forward a new class of baselines called graph-based predictors (GBPs). In contrast to the “random” baseline, GBPs leverage the known graph structure, exploiting simple graph properties to provide improved baselines against which to compare CSL methods. We discuss GBPs in general and provide a detailed study in the context of transitively closed graphs, introducing two conceptually simple baselines for this setting, the observed in-degree predictor (OIP) and the transitivity assuming predictor (TAP). While the former is straightforward to compute, for the latter we propose several simulation strategies. Moreover, we study and compare the proposed predictors theoretically, including a result showing that the OIP outperforms in expectation the “random” baseline on a subclass of latent network models featuring positive correlation among edge probabilities. Using both simulated and real biological data, we show that the proposed GBPs outperform random baselines in practice, often substantially. Some GBPs even outperform standard CSL methods (whilst being computationally cheap in practice). Our results provide a new way to assess CSL methods for interventional data.