Generating confidence intervals on biological networks.

Generating confidence intervals on biological networks.
复制标题

DOI:
10.1186/1471-2105-8-467
复制
发表时间:
2007-11-30
期刊:
影响因子:
3
通讯作者:
Stumpf MP
Stumpf MP
中科院分区:
生物学4区
文献类型:
--
作者:
Thorne T;Stumpf MP

文献摘要

参考文献

被引文献

相似文献

在网络分析中,我们经常需要某些网络统计数据的统计显着性,例如交互节点属性的相似性度量。网络的结构可能会引入节点之间的依赖性,并且通常有必要在统计分析中考虑这些依赖性。为此,我们需要某种形式的网络空模型:通常会生成网络的重新连线副本,仅保留每个节点的程度(交互数量)。我们表明,当存在可能令人困惑的附加信息时,这可能无法捕获网络结构的重要特征,并且可能导致不切实际的显着性水平。我们提出了一种新的网络重采样空模型,该模型考虑了程度序列以及可用的生物学注释。使用基因本体信息作为说明,我们展示了如何在重采样方法中解释该信息,以及这些信息对评估酿酒酵母蛋白质相互作用网络中相关性和基序丰度的统计显着性的影响。引入了 GOcardShuffle 算法,可以有效地构建网络数据的改进空模型。我们使用酿酒酵母的蛋白质相互作用网络;对于以现有数据的不同方面为条件的空模型,评估了相互作用蛋白的进化速率和表达水平之间的相关性及其统计显着性。新颖的 GOcardShuffle 方法产生了带注释的网络数据的空模型,该模型似乎更好地描述了真实生物网络的属性。一种用于生物网络数据统计分析的改进统计方法,其以可用生物信息为条件,与忽略此类注释的方法相比,会产生质量上不同的结果。特别是,我们证明网络生物组织的影响足以解释观察到的相互作用蛋白质的相似性。
In the analysis of networks we frequently require the statistical significance of some network statistic, such as measures of similarity for the properties of interacting nodes. The structure of the network may introduce dependencies among the nodes and it will in general be necessary to account for these dependencies in the statistical analysis. To this end we require some form of Null model of the network: generally rewired replicates of the network are generated which preserve only the degree (number of interactions) of each node. We show that this can fail to capture important features of network structure, and may result in unrealistic significance levels, when potentially confounding additional information is available. We present a new network resampling Null model which takes into account the degree sequence as well as available biological annotations. Using gene ontology information as an illustration we show how this information can be accounted for in the resampling approach, and the impact such information has on the assessment of statistical significance of correlations and motif-abundances in the Saccharomyces cerevisiae protein interaction network. An algorithm, GOcardShuffle, is introduced to allow for the efficient construction of an improved Null model for network data. We use the protein interaction network of S. cerevisiae; correlations between the evolutionary rates and expression levels of interacting proteins and their statistical significance were assessed for Null models which condition on different aspects of the available data. The novel GOcardShuffle approach results in a Null model for annotated network data which appears better to describe the properties of real biological networks. An improved statistical approach for the statistical analysis of biological network data, which conditions on the available biological information, leads to qualitatively different results compared to approaches which ignore such annotations. In particular we demonstrate the effects of the biological organization of the network can be sufficient to explain the observed similarity of interacting proteins.
DOI: 10.1186/1741-7007-4-39
发表时间: 2006-11-03
期刊: BMC BIOLOGY
影响因子: 5.4
作者:
de Silva, Eric;Thorne, Thomas;Ingram, Piers;Agrafioti, Ino;Swire, Jonathan;Wiuf, Carsten;Stumpf, Michael P. H.
通讯作者: Stumpf, Michael P. H.
DOI: 10.1093/nar/28.1.289
发表时间: 2000-01-01
影响因子: 14.9
作者:
Xenarios, I;Rice, DW;Eisenberg, D
通讯作者: Eisenberg, D
DOI: 10.1038/415141a
发表时间: 2002-01-10
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Bösche, M;Superti-Furga, G
通讯作者: Superti-Furga, G
DOI: 10.1073/pnas.0501179102
发表时间: 2005-03-22
影响因子: 11.1
作者:
Stumpf, MPH;Wiuf, C;May, RM
通讯作者: May, RM
DOI: 10.1093/oxfordjournals.molbev.a003913
发表时间: 2001-07-01
影响因子: 10.7
作者:
Wagner, A
通讯作者: Wagner, A