Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models

Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models
复制标题

DOI:
10.1103/physreve.87.012707
复制
发表时间:
2013-01-11
期刊:
影响因子:
2.4
通讯作者:
Aurell, Erik
Aurell, Erik
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Ekeberg, Magnus;Lovkvist, Cecilia;Aurell, Erik

文献摘要

被引文献

相似文献

蛋白质中空间位置相近的氨基酸往往会共同进化。因此,蛋白质的三维(3D)结构在进化记录中留下了相关性的回声。从这种相关性中逆向工程3D结构是结构生物学中的一个开放性问题,随着越来越多的蛋白质序列不断填充数据库,人们越来越积极地追求。在这项任务中存在一个统计推断问题,根源在于:蛋白质序列中两个位点之间的相关性可能来自第一手相互作用,但也可能通过中间位点进行网络传播;观察到的相关性不足以保证接近。将直接相互作用与间接相互作用分开是逆统计力学的一般问题的一个实例,其中的任务是从大型系统中的可观测量(磁化,相关性,样本)中学习模型参数(场,耦合)。在蛋白质序列的上下文中,该方法被称为直接耦合分析。在这里,我们表明,pseudolikestrium方法,适用于21个国家的波茨模型描述的统计特性的家庭的进化相关的蛋白质,显着优于现有的方法直接耦合分析,后者是基于标准的平均场技术。这种改进的性能还依赖于耦合强度的修改分数。使用各种蛋白质家族的特定序列实例的已知晶体结构来验证结果。实现新方法的代码可以在http://plmdca.csc.kth.se/上找到。DOI:10.1103/PhysRevE.87.012707
Spatially proximate amino acids in a protein tend to coevolve. A protein's three-dimensional (3D) structure hence leaves an echo of correlations in the evolutionary record. Reverse engineering 3D structures from such correlations is an open problem in structural biology, pursued with increasing vigor as more and more protein sequences continue to fill the data banks. Within this task lies a statistical inference problem, rooted in the following: correlation between two sites in a protein sequence can arise from firsthand interaction but can also be network-propagated via intermediate sites; observed correlation is not enough to guarantee proximity. To separate direct from indirect interactions is an instance of the general problem of inverse statistical mechanics, where the task is to learn model parameters (fields, couplings) from observables (magnetizations, correlations, samples) in large systems. In the context of protein sequences, the approach has been referred to as direct-coupling analysis. Here we show that the pseudolikelihood method, applied to 21-state Potts models describing the statistical properties of families of evolutionarily related proteins, significantly outperforms existing approaches to the direct-coupling analysis, the latter being based on standard mean-field techniques. This improved performance also relies on a modified score for the coupling strength. The results are verified using known crystal structures of specific sequence instances of various protein families. Code implementing the new method can be found at http://plmdca.csc.kth.se/. DOI: 10.1103/PhysRevE.87.012707