Integrative analysis of many weighted co-expression networks using tensor computation.

Integrative analysis of many weighted co-expression networks using tensor computation.
复制标题

DOI:
10.1371/journal.pcbi.1001106
复制
发表时间:
2011-06
影响因子:
4.3
通讯作者:
Zhou XJ
Zhou XJ
中科院分区:
生物学2区
文献类型:
--
作者:
Li W;Liu CC;Zhang T;Li H;Waterman MS;Zhou XJ

文献摘要

参考文献

被引文献

相似文献

生物网络的快速积累提出了新的挑战,需要强大的综合分析工具。大多数能够同时分析大量网络的现有方法主要是为未加权网络设计的,并且不能轻易扩展到加权网络。然而,众所周知,通过用阈值二分加权网络的边缘来将加权网络转换为未加权网络通常会导致信息丢失。我们开发了一种新颖的、基于张量的计算框架,用于在大量加权网络中挖掘循环重子图。具体来说,我们将循环重子图识别问题表述为具有稀疏约束的重 3D 子张量发现问题。我们通过设计多阶段凸松弛协议和非均匀边缘采样技术描述了解决此问题的有效方法。我们将我们的方法应用于 130 个共表达网络,并识别了 11,394 个重复重子图,分为 2,810 个家族。我们通过对大量已编译的生物知识库进行验证,证明了所识别的子图代表了有意义的生物模块。我们还表明,重子图有意义的可能性随着其在多个网络中的重复而显着增加,这凸显了生物网络分析综合方法的重要性。此外,我们基于加权图的方法可以检测到许多使用未加权图会被忽略的模式。此外,我们还发现了大量主要出现在特定表型下的模块。该分析产生了基因网络模块在全基因组范围内映射到现象组上的结果。最后,通过比较多个数据集的模块活动,我们发现了蛋白质复合物网络和转录调控网络中的高阶动态协作性。为了研究复杂的细胞网络,我们需要考虑它们在许多不同的实验或生理条件下的动态拓扑。因此,对大量生物网络的综合分析成为数据挖掘的新挑战。最近,我们和其他人提出了几种跨许多生物网络(主要关注未加权网络)的循环模式挖掘算法。然而,到目前为止,还没有专门设计的算法来挖掘大量加权大规模网络中的重复模式。在本文中,我们提出了一个计算框架来识别来自许多加权大型网络的循环重子图。通过将我们的方法应用于 130 个共表达网络,我们确定了一个很可能代表功能模块、转录模块和蛋白质复合物的模块图谱。未加权的网络分析会忽略其中许多模块。此外,许多已识别的模块构成了特定表型的特征。最后,我们证明我们的结果有助于蛋白质复合物网络和转录调控网络中高阶动态协调的研究。
The rapid accumulation of biological networks poses new challenges and calls for powerful integrative analysis tools. Most existing methods capable of simultaneously analyzing a large number of networks were primarily designed for unweighted networks, and cannot easily be extended to weighted networks. However, it is known that transforming weighted into unweighted networks by dichotomizing the edges of weighted networks with a threshold generally leads to information loss. We have developed a novel, tensor-based computational framework for mining recurrent heavy subgraphs in a large set of massive weighted networks. Specifically, we formulate the recurrent heavy subgraph identification problem as a heavy 3D subtensor discovery problem with sparse constraints. We describe an effective approach to solving this problem by designing a multi-stage, convex relaxation protocol, and a non-uniform edge sampling technique. We applied our method to 130 co-expression networks, and identified 11,394 recurrent heavy subgraphs, grouped into 2,810 families. We demonstrated that the identified subgraphs represent meaningful biological modules by validating against a large set of compiled biological knowledge bases. We also showed that the likelihood for a heavy subgraph to be meaningful increases significantly with its recurrence in multiple networks, highlighting the importance of the integrative approach to biological network analysis. Moreover, our approach based on weighted graphs detects many patterns that would be overlooked using unweighted graphs. In addition, we identified a large number of modules that occur predominately under specific phenotypes. This analysis resulted in a genome-wide mapping of gene network modules onto the phenome. Finally, by comparing module activities across many datasets, we discovered high-order dynamic cooperativeness in protein complex networks and transcriptional regulatory networks. To study complex cellular networks, we need to consider their dynamic topologies under many different experimental or physiological conditions. Integrative analysis over large numbers of massive biological networks thus emerges as a new challenge in data mining. Recently, we and others have proposed several algorithms for recurrent pattern mining across many () biological networks (with the main focus on unweighted networks). However, thus far no algorithms have been specifically designed to mine recurrent patterns across a large collection of weighted massive networks. In this paper, we propose a computational framework to identify recurrent heavy subgraphs from many weighted large networks. By applying our method to 130 co-expression networks, we identified an atlas of modules that are highly likely to represent functional modules, transcriptional modules, and protein complexes. Many of these modules would be overlooked with unweighted networks analysis. Furthermore, many of the identified modules constituted signatures of specific phenotypes. Finally, we demonstrated that our results facilitate the study of high-order dynamic coordination in protein complex networks and transcriptional regulatory networks.
DOI: 10.1037/h0054245
发表时间: 1952-01-01
影响因子: 22.4
作者:
CATTELL, RB
通讯作者: CATTELL, RB
DOI: 10.1089/cmb.2009.0099
发表时间: 2009-08-01
影响因子: 1.7
作者:
Flannick, Jason;Novak, Antal;Batzoglou, Serafim
通讯作者: Batzoglou, Serafim
DOI: 10.1073/pnas.97.18.10101
发表时间: 2000-08-29
影响因子: 11.1
作者:
Alter, O;Brown, PO;Botstein, D
通讯作者: Botstein, D
DOI: 10.1093/nar/gkm1001
发表时间: 2008-01
影响因子: 14.9
作者:
Breitkreutz BJ;Stark C;Reguly T;Boucher L;Breitkreutz A;Livstone M;Oughtred R;Lackner DH;Bähler J;Wood V;Dolinski K;Tyers M
通讯作者: Tyers M
DOI: 10.1093/bioinformatics/btm222
发表时间: 2007-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Huang, Yu;Li, Haifeng;Zhou, Xianghong Jasmine
通讯作者: Zhou, Xianghong Jasmine