A correlated motif approach for finding short linear motifs from protein interaction networks.

A correlated motif approach for finding short linear motifs from protein interaction networks.
复制标题

DOI:
10.1186/1471-2105-7-502
复制
发表时间:
2006-11-16
期刊:
影响因子:
3
通讯作者:
Ng SK
Ng SK
中科院分区:
生物学4区
文献类型:
--
作者:
Tan SH;Hugo W;Sung WK;Ng SK

文献摘要

参考文献

被引文献

相似文献

生物回路和疾病通路的一类重要的相互作用开关是短结合基序。然而,寻找这些结合基序的生物学实验往往是费力和昂贵的。随着蛋白质相互作用数据的可用性,可以通过计算发现新的结合基序:通过将标准基序提取算法应用于蛋白质序列集,每个蛋白质序列集与共同蛋白质或具有相似性质的蛋白质组相互作用。潜在的假设是,具有共同相互作用伴侣的蛋白质将共享一些共同的结合基序。虽然已经用这种方法发现了新的结合基序,但是如果蛋白质与非常少的其他蛋白质相互作用或者当蛋白质组的先验知识不可用或错误时,则不适用。输入交互数据中的实验噪声会进一步恶化这种方法的糟糕性能。我们提出了一种新的方法,发现相关的短序列模体的蛋白质-蛋白质相互作用的数据,以有效地规避上述限制。相关基序是指那些始终只在相互作用的蛋白质序列对中共同出现的基序,并且可能直接或间接地相互作用以介导相互作用。我们采用了(l,d)-模体模型,并制定相关的模体作为(l,d)-模体对发现问题。我们提出了一个精确的算法,D-MOTIF,以及它的近似算法,D-STAR来解决这个问题。对大量模拟数据的评估表明,我们的方法不仅消除了对任何先前蛋白质分组的需要,而且在从嘈杂的相互作用数据中提取基序时也更鲁棒。在两个生物数据集(SH 3相互作用网络和TGFβ信号网络)上的应用表明,该方法可以提取出与实际相互作用网络相对应的相关模体。本文提出的相关模体方法能够从稀疏和噪声的相互作用数据中找到相关的线性模体。这反过来又将加快新的线性结合基序的发现,并促进由它们介导的生物途径的研究。
An important class of interaction switches for biological circuits and disease pathways are short binding motifs. However, the biological experiments to find these binding motifs are often laborious and expensive. With the availability of protein interaction data, novel binding motifs can be discovered computationally: by applying standard motif extracting algorithms on protein sequence sets each interacting with either a common protein or a protein group with similar properties. The underlying assumption is that proteins with common interacting partners will share some common binding motifs. Although novel binding motifs have been discovered with such approach, it is not applicable if a protein interacts with very few other proteins or when prior knowledge of protein group is not available or erroneous. Experimental noise in input interaction data can further deteriorate the dismal performance of such approaches. We propose a novel approach of finding correlated short sequence motifs from protein-protein interaction data to effectively circumvent the above-mentioned limitations. Correlated motifs are those motifs that consistently co-occur only in pairs of interacting protein sequences, and could possibly interact with each other directly or indirectly to mediate interactions. We adopted the (l, d)-motif model and formulate finding the correlated motifs as an (l, d)-motif pair finding problem. We present both an exact algorithm, D-MOTIF, as well as its approximation algorithm, D-STAR to solve this problem. Evaluation on extensive simulated data showed that our approach not only eliminated the need for any prior protein grouping, but is also more robust in extracting motifs from noisy interaction data. Application on two biological datasets (SH3 interaction network and TGFβ signaling network) demonstrates that the approach can extract correlated motifs that correspond to actual interacting subsequences. The correlated motif approach outlined in this paper is able to find correlated linear motifs from sparse and noisy interaction data. This, in turn, will expedite the discovery of novel linear binding motifs, and facilitate the studies of biological pathways mediated by them.
DOI: 10.1371/journal.pbio.0030405
发表时间: 2005-12
期刊: PLoS biology
影响因子: 9.8
作者:
Neduva V;Linding R;Su-Angrand I;Stark A;de Masi F;Gibson TJ;Lewis J;Serrano L;Russell RB
通讯作者: Russell RB
DOI: 10.1093/bioinformatics/14.1.55
发表时间: 1998-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Rigoutsos, I;Floratos, A
通讯作者: Floratos, A
DOI: 10.1093/bioinformatics/btg118
发表时间: 2003-05-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Ng, SK;Zhang, Z;Tan, SH
通讯作者: Tan, SH
DOI: 10.1101/gr.153002
发表时间: 2002-10-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Deng, MH;Mehta, S;Chen, T
通讯作者: Chen, T
DOI: 10.1002/pmic.200300632
发表时间: 2004-03-01
期刊: PROTEOMICS
影响因子: 3.4
作者:
Hu, H;Columbus, J;Herrero, JJ
通讯作者: Herrero, JJ