Protein-protein interactions more conserved within species than across species

Protein-protein interactions more conserved within species than across species
复制标题

DOI:
10.1371/journal.pcbi.0020079
复制
发表时间:
2006-07-01
影响因子:
4.3
通讯作者:
Rost, Burkhard
Rost, Burkhard
中科院分区:
生物学2区
文献类型:
--
作者:
Mika, Sven;Rost, Burkhard

文献摘要

被引文献

相似文献

蛋白质-蛋白质相互作用的实验性高通量研究开始为综合计算研究提供足够的数据。如今,大约有十个大型数据集,每个数据集都有数千个相互作用对,粗略地采样了果蝇、人类、蠕虫和酵母中的相互作用。通过更仔细、更详细的生化实验,还鉴定出了另外大约 55,000 对相互作用的蛋白质。大多数相互作用都是在原核生物和简单真核生物中通过实验观察到的。在哺乳动物等高等真核生物中观察到的相互作用很少。人们普遍认为,哺乳动物中的途径可以通过与模型生物(例如动物)的同源性来推断。 g。将两种酵母蛋白相互作用的实验观察结果转移到人类中相应的两种蛋白质也相互作用。相互作用保守的两对通常被描述为间同源物。这项研究的目标是对此类推论进行大规模综合分析,即对内同源物的进化保守性进行分析。在这里,我们引入了一种新的评分来测量蛋白质-蛋白质相互作用数据集之间的重叠。这一衡量标准似乎反映了数据的整体质量,并且是我们大规模分析中得出的两个令人惊讶的结果的基础。首先,基于同源性的物理蛋白质-蛋白质相互作用的推断似乎远没有预期那么成功。事实上,只有在序列相似性极高的情况下,这样的推论才是准确的。其次,最令人惊讶的是,通过序列相似性识别相互作用伙伴对于同一生物体内的蛋白质对比物种之间的蛋白质对更可靠。我们的分析强调,即使在同一生物体上使用相同类型的实验,不同数据集之间的差异也很大。这一现实极大地限制了基于同源性的交互转移的能力。特别是,对遥远模型生物中相互作用的实验探测必须谨慎进行。更全面的蛋白质-蛋白质网络图像需要结合许多高通量方法,包括计算机推断和预测。
Experimental high-throughput studies of protein-protein interactions are beginning to provide enough data for comprehensive computational studies. Today, about ten large data sets, each with thousands of interacting pairs, coarsely sample the interactions in fly, human, worm, and yeast. Another about 55,000 pairs of interacting proteins have been identified by more careful, detailed biochemical experiments. Most interactions are experimentally observed in prokaryotes and simple eukaryotes; very few interactions are observed in higher eukaryotes such as mammals. It is commonly assumed that pathways in mammals can be inferred through homology to model organisms, e. g. the experimental observation that two yeast proteins interact is transferred to infer that the two corresponding proteins in human also interact. Two pairs for which the interaction is conserved are often described as interologs. The goal of this investigation was a large-scale comprehensive analysis of such inferences, i.e. of the evolutionary conservation of interologs. Here, we introduced a novel score for measuring the overlap between protein-protein interaction data sets. This measure appeared to reflect the overall quality of the data and was the basis for our two surprising results from our large-scale analysis. Firstly, homology-based inferences of physical protein-protein interactions appeared far less successful than expected. In fact, such inferences were accurate only for extremely high levels of sequence similarity. Secondly, and most surprisingly, the identification of interacting partners through sequence similarity was significantly more reliable for protein pairs within the same organism than for pairs between species. Our analysis underlined that the discrepancies between different datasets are large, even when using the same type of experiment on the same organism. This reality considerably constrains the power of homology-based transfer of interactions. In particular, the experimental probing of interactions in distant model organisms has to be undertaken with some caution. More comprehensive images of protein-protein networks will require the combination of many high-throughput methods, including in silico inferences and predictions.