Assessing statistical significance in causal graphs

Assessing statistical significance in causal graphs
复制标题

DOI:
10.1186/1471-2105-13-35
复制
发表时间:
2012-02-20
期刊:
影响因子:
3
通讯作者:
Ziemek, Daniel
Ziemek, Daniel
中科院分区:
生物学4区
文献类型:
--
作者:
Chindelevitch, Leonid;Loh, Po-Ru;Ziemek, Daniel

文献摘要

被引文献

相似文献

背景:因果图是分析生物学数据集的一种日益流行的工具。具体地说,带符号的因果图--有向图的边另外有一个符号表示上调或下调--可以用来对细胞内的监管网络进行建模。这样的模型可以预测生物实体调控的下游影响;相反,它们也能够推断观察到的表达变化背后的致病因子。然而,由于其复杂性,带符号的因果图模型在评估统计显著性方面提出了特殊的挑战。在这篇文章中,我们框架并解决了在计算适当的零分布进行假设检验时出现的两个基本计算问题。结果:首先,我们展示了如何计算上调、下调或两者都不上调的基因转录本的观察分类和模型预测分类之间的一致性的p值。具体地说,在观察到的分类的零分布被随机化的情况下,分类达到相同程度的可能性有多大?这个问题因其数学形式而被称为“三元点积分布”,可以看作是Fisher对三元变量的精确检验的推广。我们提出了两种计算三值点乘积分布的高效算法,并从解析和数值两个方面研究了它的组合结构,建立了计算复杂性的界。其次,我们提出了一种对因果图进行高效随机抽样的算法。这允许在不同的、同样重要的零分布下进行p值计算,该分布是通过随机化图形拓扑但保持其基本结构不变而获得的:连通性以及每个顶点的正和负进出度。我们给出了一个从该分布中随机抽样图的算法。我们还强调了有符号因果图特有的理论挑战;之前关于图随机化的工作已经研究了无向图和有向但无符号图。结论:我们给出了应用因果图方法这一生物网络分析的强大工具所必需的两个统计意义问题的算法解决方案。我们提出的算法既快速又被证明是正确的。我们的工作可能在非生物背景下也有独立的兴趣,因为它推广了在其他领域已被广泛研究的数学结果。
Background: Causal graphs are an increasingly popular tool for the analysis of biological datasets. In particular, signed causal graphs-directed graphs whose edges additionally have a sign denoting upregulation or downregulation-can be used to model regulatory networks within a cell. Such models allow prediction of downstream effects of regulation of biological entities; conversely, they also enable inference of causative agents behind observed expression changes. However, due to their complex nature, signed causal graph models present special challenges with respect to assessing statistical significance. In this paper we frame and solve two fundamental computational problems that arise in practice when computing appropriate null distributions for hypothesis testing.Results: First, we show how to compute a p-value for agreement between observed and model-predicted classifications of gene transcripts as upregulated, downregulated, or neither. Specifically, how likely are the classifications to agree to the same extent under the null distribution of the observed classification being randomized? This problem, which we call "Ternary Dot Product Distribution" owing to its mathematical form, can be viewed as a generalization of Fisher's exact test to ternary variables. We present two computationally efficient algorithms for computing the Ternary Dot Product Distribution and investigate its combinatorial structure analytically and numerically to establish computational complexity bounds. Second, we develop an algorithm for efficiently performing random sampling of causal graphs. This enables p-value computation under a different, equally important null distribution obtained by randomizing the graph topology but keeping fixed its basic structure: connectedness and the positive and negative in-and out-degrees of each vertex. We provide an algorithm for sampling a graph from this distribution uniformly at random. We also highlight theoretical challenges unique to signed causal graphs; previous work on graph randomization has studied undirected graphs and directed but unsigned graphs.Conclusion: We present algorithmic solutions to two statistical significance questions necessary to apply the causal graph methodology, a powerful tool for biological network analysis. The algorithms we present are both fast and provably correct. Our work may be of independent interest in non-biological contexts as well, as it generalizes mathematical results that have been studied extensively in other fields.