Inter-Protein Sequence Co-Evolution Predicts Known Physical Interactions in Bacterial Ribosomes and the Trp Operon.

Inter-Protein Sequence Co-Evolution Predicts Known Physical Interactions in Bacterial Ribosomes and the Trp Operon.
复制标题

DOI:
10.1371/journal.pone.0149166
复制
发表时间:
2016
期刊:
影响因子:
3.7
通讯作者:
Pagnani A
Pagnani A
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Feinauer C;Szurmant H;Weigt M;Pagnani A

文献摘要

被引文献

相似文献

蛋白质之间的相互作用是构成几乎所有生物过程的基本机制。许多重要的相互作用在各种物种中都是保守的。保持相互作用的需要导致伙伴蛋白之间界面上的残基之间高度协同进化。从快速增长的序列数据库中推断蛋白质-蛋白质相互作用网络是当今系统生物学中最艰巨的任务之一。本文提出了一种基于蛋白质残基对间协同进化的直接耦合分析的新方法。我们使用核糖体和Trp操纵子蛋白作为测试案例:对于小的分别。大核糖体亚基我们的方法预测蛋白质相互作用伙伴的真阳性率分别为70%。在前10个预测中占90%,面积分别为0.69。所有预测的ROC曲线下的0.81。在Trp操纵子中,它将两个最大的相互作用分数分配给仅有的两个实验上已知的相互作用。在残基相互作用的水平上,我们表明,对于小的和大的核糖体亚基,我们的方法预测了系统中相互作用的残基,在前20个预测中的真阳性率分别为60%和85%。我们使用人工数据来表明我们的方法的性能关键取决于联合多序列比对的大小,并分析了如果序列是从我们用于预测的同一模型中采样的,那么完美预测需要多少序列。考虑到我们方法在测试数据上的性能,我们推测它可以用于检测新的交互作用,特别是在可用序列数据快速增长的情况下。
Interaction between proteins is a fundamental mechanism that underlies virtually all biological processes. Many important interactions are conserved across a large variety of species. The need to maintain interaction leads to a high degree of co-evolution between residues in the interface between partner proteins. The inference of protein-protein interaction networks from the rapidly growing sequence databases is one of the most formidable tasks in systems biology today. We propose here a novel approach based on the Direct-Coupling Analysis of the co-evolution between inter-protein residue pairs. We use ribosomal and trp operon proteins as test cases: For the small resp. large ribosomal subunit our approach predicts protein-interaction partners at a true-positive rate of 70% resp. 90% within the first 10 predictions, with areas of 0.69 resp. 0.81 under the ROC curves for all predictions. In the trp operon, it assigns the two largest interaction scores to the only two interactions experimentally known. On the level of residue interactions we show that for both the small and the large ribosomal subunit our approach predicts interacting residues in the system with a true positive rate of 60% and 85% in the first 20 predictions. We use artificial data to show that the performance of our approach depends crucially on the size of the joint multiple sequence alignments and analyze how many sequences would be necessary for a perfect prediction if the sequences were sampled from the same model that we use for prediction. Given the performance of our approach on the test data we speculate that it can be used to detect new interactions, especially in the light of the rapid growth of available sequence data.