Direct-coupling analysis of residue coevolution captures native contacts across many protein families

Direct-coupling analysis of residue coevolution captures native contacts across many protein families
复制标题

DOI:
10.1073/pnas.1111471108
复制
发表时间:
2011-12-06
影响因子:
11.1
通讯作者:
Weigt, Martin
Weigt, Martin
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Morcos, Faruck;Pagnani, Andrea;Weigt, Martin

文献摘要

被引文献

相似文献

同源蛋白在三维结构上的相似性对其序列变异性施加了很强的限制。长期以来,人们一直认为,不同序列位置的氨基酸组成之间的相关性可以用来推断三级蛋白质结构内的空间联系。对这种推断至关重要的是分解直接和间接相关性的能力,正如最近引入的直接耦合分析(DCA)所完成的那样。在这里,我们开发了一个计算效率高的DCA实现,它允许我们评估DCA对大量蛋白质结构域的接触预测的准确性,纯粹基于序列信息。DCA显示出大量正确预测的接触,概括了所检查的大多数蛋白质结构域的接触图的全局结构。此外,我们的分析捕获了结构域内残基接触之外的清晰信号,例如,由替代蛋白质构象、配体介导的残基偶联和蛋白质低聚物中的结构域间相互作用引起的。我们的研究结果表明,DCA预测的接触可以作为可靠的指导,以促进替代蛋白质构象的计算预测,蛋白质复合物的形成,甚至蛋白质结构域的从头预测,这取决于大量同源序列的存在,这些序列由于基因组测序的进步而迅速可用。
The similarity in the three-dimensional structures of homologous proteins imposes strong constraints on their sequence variability. It has long been suggested that the resulting correlations among amino acid compositions at different sequence positions can be exploited to infer spatial contacts within the tertiary protein structure. Crucial to this inference is the ability to disentangle direct and indirect correlations, as accomplished by the recently introduced direct-coupling analysis (DCA). Here we develop a computationally efficient implementation of DCA, which allows us to evaluate the accuracy of contact prediction by DCA for a large number of protein domains, based purely on sequence information. DCA is shown to yield a large number of correctly predicted contacts, recapitulating the global structure of the contact map for the majority of the protein domains examined. Furthermore, our analysis captures clear signals beyond intradomain residue contacts, arising, e.g., from alternative protein conformations, ligand-mediated residue couplings, and interdomain interactions in protein oligomers. Our findings suggest that contacts predicted by DCA can be used as a reliable guide to facilitate computational predictions of alternative protein conformations, protein complex formation, and even the de novo prediction of protein domain structures, contingent on the existence of a large number of homologous sequences which are being rapidly made available due to advances in genome sequencing.