Predicting and validating protein interactions using network structure.

Predicting and validating protein interactions using network structure.
复制标题

DOI:
10.1371/journal.pcbi.1000118
复制
发表时间:
2008-07-25
影响因子:
4.3
通讯作者:
Reinert G
Reinert G
中科院分区:
生物学2区
文献类型:
--
作者:
Chen PY;Deane CM;Reinert G

文献摘要

参考文献

被引文献

相似文献

蛋白质相互作用在细胞功能中起着至关重要的作用。由于用于检测和验证蛋白质相互作用的实验技术是耗时的,因此需要用于该任务的计算方法。蛋白质相互作用似乎形成了一个网络,具有相对较高的局部聚类程度。在本文中,我们利用这种聚类提出了一个分数的基础上观察到的蛋白质相互作用的三联体。该评分利用蛋白质特征和网络特性。我们的评分基于三胞胎,以补充现有的技术预测蛋白质相互作用,优于他们的数据集显示高度的聚类。预测的相互作用得分很高,对测试措施的准确性。与仅来自成对相互作用的类似评分相比,三联体评分显示出更高的灵敏度和特异性。通过查看具体的例子,我们展示了如何丰富和验证一组实验性的相互作用。作为这项工作的一部分,我们还研究了不同的先验数据库对预测准确性的影响,并发现来自同一王国的相互作用比来自不同王国的相互作用给出了更好的结果,这表明网络之间可能存在根本性的差异。这些结果都强调了网络结构的重要性,有助于准确预测蛋白质相互作用。蛋白质相互作用数据集和我们分析中使用的程序,以及预测和验证列表,可在http://www.stats.ox.ac.uk/bioinfo/resources/PredictingInteractions上获得。为了理解生物体内的复杂活动,生物体内发生的蛋白质相互作用的完整且无错误的网络将是向前迈出的重要一步。大量的实验数据为我们研究蛋白质相互作用的复杂行为提供了机会。然而,由于数据集中的高假阳性和假阴性率,这些研究的力量受到限制。我们提出了一种基于网络的方法,利用蛋白质相互作用网络中的聚类趋势,来验证实验数据和预测未知的相互作用。多种蛋白质特征的整合(即,结构、功能等)允许我们的预测方法显着优于其他两种方法的基础上的同源性和蛋白质结构域的关系数据集,其中包含大量的相互作用,但没有太多的详细信息的蛋白质参与的相互作用。此外,我们基于三元交互模式的预测评分比配对方法有所提高,这表明网络结构的重要性。此外,使用池的相互作用作为先验信息,我们发现的证据,真核生物和原核生物之间的蛋白质相互作用网络的根本差异。
Protein interactions play a vital part in the function of a cell. As experimental techniques for detection and validation of protein interactions are time consuming, there is a need for computational methods for this task. Protein interactions appear to form a network with a relatively high degree of local clustering. In this paper we exploit this clustering by suggesting a score based on triplets of observed protein interactions. The score utilises both protein characteristics and network properties. Our score based on triplets is shown to complement existing techniques for predicting protein interactions, outperforming them on data sets which display a high degree of clustering. The predicted interactions score highly against test measures for accuracy. Compared to a similar score derived from pairwise interactions only, the triplet score displays higher sensitivity and specificity. By looking at specific examples, we show how an experimental set of interactions can be enriched and validated. As part of this work we also examine the effect of different prior databases upon the accuracy of prediction and find that the interactions from the same kingdom give better results than from across kingdoms, suggesting that there may be fundamental differences between the networks. These results all emphasize that network structure is important and helps in the accurate prediction of protein interactions. The protein interaction data set and the program used in our analysis, and a list of predictions and validations, are available at http://www.stats.ox.ac.uk/bioinfo/resources/PredictingInteractions. For understanding the complex activities within an organism, a complete and error-free network of protein interactions which occur in the organism would be a significant step forward. The large amount of experimentally derived data now available has provided us with a chance to study the complicated behaviour of protein interactions. The power of such studies, however, has been limited due to the high false positive and false negative rates in the datasets. We propose a network-based method, taking advantage of the tendency of clustering in protein interaction networks, to validate experimental data and to predict unknown interactions. The integration of multiple protein characteristics (i.e., structure, function, etc.) allows our predictive method to significantly outperform two other approaches based on homology and protein-domain relationships on datasets which contain a large amount of interactions, but not much detailed information on the proteins involved in the interactions. In addition, our predictive score based on triadic interaction patterns improves over a pair-wise approach, suggesting the importance of network structure. Moreover, using pooled interactions as prior information, we find evidence for fundamental differences in protein interaction networks between eukaryotes and prokaryotes.
通过同源性生成的网络的聚类分析:自动鉴定参与癌症转移的重要蛋白质群落。
DOI: 10.1186/1471-2105-7-2
发表时间: 2006-01-06
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Jonsson, PF;Cavanna, T;Zicha, D;Bates, PA
通讯作者: Bates, PA
DOI: 10.1186/gb-2006-7-11-120
发表时间: 2006
期刊: Genome biology
影响因子: 12.3
作者:
Hart GT;Ramani AK;Marcotte EM
通讯作者: Marcotte EM
DOI: 10.1093/nar/30.1.268
发表时间: 2002-01-01
影响因子: 14.9
作者:
Gough, J;Chothia, C
通讯作者: Chothia, C
DOI: 10.1371/journal.pcbi.0020079
发表时间: 2006-07-01
影响因子: 4.3
作者:
Mika, Sven;Rost, Burkhard
通讯作者: Rost, Burkhard
DOI: 10.1038/35075138
发表时间: 2001-05-03
期刊: NATURE
影响因子: 64.8
作者:
Jeong, H;Mason, SP;Oltvai, ZN
通讯作者: Oltvai, ZN