Fast pseudolikelihood maximization for direct-coupling analysis of protein structure from many homologous amino-acid sequences

Fast pseudolikelihood maximization for direct-coupling analysis of protein structure from many homologous amino-acid sequences
复制标题

DOI:
10.1016/j.jcp.2014.07.024
复制
发表时间:
2014-11-01
影响因子:
4.1
通讯作者:
Aurell, Erik
Aurell, Erik
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
Ekeberg, Magnus;Hartonen, Tuomo;Aurell, Erik

文献摘要

被引文献

相似文献

直接偶联分析是一组通过从数据中学习指数家族中的生成模型来获取蛋白质家族中共进化残基信息的方法。在实际大小的蛋白质家族中,这种学习只能近似完成,并且在推理精度和计算速度之间存在权衡。我们在这里表明,早期引入的l(2)-正则化的伪概率最大化方法称为plmDCA可以修改为易于并行化,以及在单个处理器上固有地更快,在精度上的差异可以忽略不计。我们测试的新化身的方法对143蛋白质家族/结构对从蛋白质家族数据库(PFAM),这类算法的最大的测试之一。(C)2014爱思唯尔公司保留所有权利。
Direct-coupling analysis is a group of methods to harvest information about coevolving residues in a protein family by learning a generative model in an exponential family from data. In protein families of realistic size, this learning can only be done approximately, and there is a trade-off between inference precision and computational speed. We here show that an earlier introduced l(2)-regularized pseudolikelihood maximization method called plmDCA can be modified as to be easily parallelizable, as well as inherently faster on a single processor, at negligible difference in accuracy. We test the new incarnation of the method on 143 protein family/structure-pairs from the Protein Families database (PFAM), one of the larger tests of this class of algorithms to date. (C) 2014 Elsevier Inc. Allrightsreserved.