A large-scale conformation sampling and evaluation server for protein tertiary structure prediction and its assessment in CASP11.

A large-scale conformation sampling and evaluation server for protein tertiary structure prediction and its assessment in CASP11.
复制标题

DOI:
10.1186/s12859-015-0775-x
复制
发表时间:
2015-10-23
期刊:
影响因子:
3
通讯作者:
Cheng J
Cheng J
中科院分区:
生物学4区
文献类型:
--
作者:
Li J;Cao R;Cheng J

文献摘要

被引文献

相似文献

随着基因组时代产生越来越多的蛋白质序列,从序列预测蛋白质结构对于阐明这些蛋白质的分子细节和功能以进行生物医学研究变得非常重要。传统的基于模板的蛋白质结构预测方法往往侧重于识别最佳模板、生成最佳比对以及应用最佳能量函数对模型进行排序,但由于难以获得最佳模板、比对和模型,往往无法达到最佳性能。我们开发了大规模构象采样和评估方法及其服务器,以提高蛋白质结构预测的可靠性和鲁棒性。第一步,我们的方法使用各种比对方法对相关和互补模板进行采样,并生成替代和多样化的目标模板比对,使用模板和比对组合协议来组合比对,并使用基于模板和无模板的建模方法来生成目标蛋白的构象池。第二步,使用大量的蛋白质模型质量评估方法对蛋白质模型池中的模型进行评估和排名,并结合异常处理策略来处理模型排名中任何额外的失败。该方法被实现为两个蛋白质结构预测服务器:MULTICOM-CONSTRUCT和MULTICOM-CLUSTER,参加了2014年第11届蛋白质结构预测技术关键评估(CASP11)。这两个服务器被评为最佳10个服务器预测器之一。我们的服务器在CASP11中的良好性能证明了大规模构象采样和评估的有效性和鲁棒性。 MULTICOM 服务器位于:http://sysbio.rnet.missouri.edu/multicom_cluster/。本文的在线版本 (doi:10.1186/s12859-015-0775-x) 包含补充材料,可供授权用户使用。
With more and more protein sequences produced in the genomic era, predicting protein structures from sequences becomes very important for elucidating the molecular details and functions of these proteins for biomedical research. Traditional template-based protein structure prediction methods tend to focus on identifying the best templates, generating the best alignments, and applying the best energy function to rank models, which often cannot achieve the best performance because of the difficulty of obtaining best templates, alignments, and models. We developed a large-scale conformation sampling and evaluation method and its servers to improve the reliability and robustness of protein structure prediction. In the first step, our method used a variety of alignment methods to sample relevant and complementary templates and to generate alternative and diverse target-template alignments, used a template and alignment combination protocol to combine alignments, and used template-based and template-free modeling methods to generate a pool of conformations for a target protein. In the second step, it used a large number of protein model quality assessment methods to evaluate and rank the models in the protein model pool, in conjunction with an exception handling strategy to deal with any additional failure in model ranking. The method was implemented as two protein structure prediction servers: MULTICOM-CONSTRUCT and MULTICOM-CLUSTER that participated in the 11th Critical Assessment of Techniques for Protein Structure Prediction (CASP11) in 2014. The two servers were ranked among the best 10 server predictors. The good performance of our servers in CASP11 demonstrates the effectiveness and robustness of the large-scale conformation sampling and evaluation. The MULTICOM server is available at: http://sysbio.rnet.missouri.edu/multicom_cluster/. The online version of this article (doi:10.1186/s12859-015-0775-x) contains supplementary material, which is available to authorized users.