A max-flow based approach to the identification of protein complexes using protein interaction and microarray data.

A max-flow based approach to the identification of protein complexes using protein interaction and microarray data.
复制标题

DOI:
10.1142/9781848162648_0005
复制
发表时间:
2008
期刊:
Computational systems bioinformatics. Computational Systems Bioinformatics Conference
影响因子:
--
通讯作者:
Jianxing Feng;Rui Jiang;Tao Jiang
Jianxing Feng;Rui Jiang;Tao Jiang
中科院分区:
其他
文献类型:
--
作者:
Jianxing Feng;Rui Jiang;Tao Jiang

文献摘要

被引文献

相似文献

高通量技术的出现产生了丰富的蛋白质相互作用(PPI)数据和微阵列基因表达谱,为利用计算方法鉴定新型蛋白质复合物提供了巨大的机会。虽然它已被证明在文献中,单独使用蛋白质-蛋白质相互作用数据的方法可以成功地预测大量的蛋白质复合物,基因表达谱的掺入可以帮助完善推定的复合物,从而提高计算方法的准确性。结合蛋白质相互作用数据和微阵列基因表达谱,我们提出了一种新的图片段化算法(GFA)的蛋白质复合物识别。GFA是从经典的最大流算法中找到(加权)稠密子图,首先在蛋白质-蛋白质相互作用网络中找到大的(加权)稠密子图,然后通过在微阵列数据中相应的对数倍数变化来适当地加权其节点,将每个这样的子图迭代地分解成片段,直到片段子图足够小。我们对三个广泛使用的蛋白质-蛋白质相互作用数据集进行了广泛的测试,并与最新的蛋白质复合物鉴定方法进行了比较,证明了我们的方法在准确性,效率和预测新型蛋白质复合物的能力方面具有上级性能。鉴于我们的方法已经达到的高特异性(或精确度),我们推测我们的预测结果意味着超过200种新的蛋白质复合物。
The emergence of high-throughput technologies leads to abundant protein-protein interaction (PPI) data and microarray gene expression profiles, and provides a great opportunity for the identification of novel protein complexes using computational methods. Although it has been demonstrated in the literature that methods using protein-protein interaction data alone can successfully predict a large number of protein complexes, the incorporation of gene expression profiles could help refine the putative complexes and hence improve the accuracy of the computational methods. By combining protein-protein interaction data and microarray gene expression profiles, we propose a novel Graph Fragmentation Algorithm (GFA) for protein complex identification. Adapted from a classical max-flow algorithm for finding the (weighted) densest subgraphs, GFA first finds large (weighted) dense subgraphs in a protein-protein interaction network and then breaks each such subgraph into fragments iteratively by weighting its nodes appropriately in terms of their corresponding log fold changes in the microarray data, until the fragment subgraphs are sufficiently small. Our extensive tests on three widely used protein-protein interaction datasets and comparisons with the latest methods for protein complex identification demonstrate the superior performance of our method in terms of accuracy, efficiency, and capability in predicting novel protein complexes. Given the high specificity (or precision) that our method has achieved, we conjecture that our prediction results imply more than 200 novel protein complexes.