Improving predicted protein loop structure ranking using a Pareto-optimality consensus method.

Improving predicted protein loop structure ranking using a Pareto-optimality consensus method.
复制标题

DOI:
10.1186/1472-6807-10-22
复制
发表时间:
2010-07-20
影响因子:
--
通讯作者:
Jakobsson E
Jakobsson E
中科院分区:
生物4区
文献类型:
--
作者:
Li Y;Rata I;Chiu SW;Jakobsson E

文献摘要

参考文献

被引文献

相似文献

精确的蛋白质环结构模型对于理解许多蛋白质的功能是非常重要的。通过将天然或近天然模型与错误折叠的模型区分开来来识别天然或近天然模型是蛋白质环结构预测的关键步骤。我们已经开发了一个帕累托最优共识(POC)的方法,这是一个共识模型的排名方法,整合多个知识或物理为基础的评分功能。识别模型集中最佳质量的模型的过程包括:1)识别相对于一组评分函数位于帕累托最优前沿的模型,以及2)基于与其余模型的模糊优势关系对它们进行排序。我们将POC方法应用于大量的诱饵集循环的4- 12-残基长度使用的功能空间组成的几个精心选择的评分功能:Rosetta,DOPE,DDFIRE,OPLS-AA,和一个三重骨架二面角的潜力,在我们的实验室开发。我们的计算结果表明,帕累托最优诱饵集通常由集合中总诱饵的约20%或更少组成,并且对超过99%的环路目标中的最佳或接近最佳诱饵具有良好的覆盖。与在诱饵集合中产生最佳选择准确度的个体评分函数相比,POC方法在区分天然构象、将近天然模型(与天然的RMSD < 0.5A)识别为排名靠前的以及在排名靠前的5个模型中选择至少一个近天然模型方面分别产生23%、37%和64%的假阳性。在来自膜蛋白环的诱饵集合中也发现了POC方法的类似有效性。此外,POC方法在模型排名中优于其他常用的共识策略,例如按编号排名,按排名排名,按投票排名和基于回归的方法。通过集成基于帕累托最优性和模糊优势的多个基于知识和物理的评分函数,POC方法在区分最佳回路模型与回路模型集合内的其他回路模型方面是有效的。
Accurate protein loop structure models are important to understand functions of many proteins. Identifying the native or near-native models by distinguishing them from the misfolded ones is a critical step in protein loop structure prediction. We have developed a Pareto Optimal Consensus (POC) method, which is a consensus model ranking approach to integrate multiple knowledge- or physics-based scoring functions. The procedure of identifying the models of best quality in a model set includes: 1) identifying the models at the Pareto optimal front with respect to a set of scoring functions, and 2) ranking them based on the fuzzy dominance relationship to the rest of the models. We apply the POC method to a large number of decoy sets for loops of 4- to 12-residue in length using a functional space composed of several carefully-selected scoring functions: Rosetta, DOPE, DDFIRE, OPLS-AA, and a triplet backbone dihedral potential developed in our lab. Our computational results show that the sets of Pareto-optimal decoys, which are typically composed of ~20% or less of the overall decoys in a set, have a good coverage of the best or near-best decoys in more than 99% of the loop targets. Compared to the individual scoring function yielding best selection accuracy in the decoy sets, the POC method yields 23%, 37%, and 64% less false positives in distinguishing the native conformation, indentifying a near-native model (RMSD < 0.5A from the native) as top-ranked, and selecting at least one near-native model in the top-5-ranked models, respectively. Similar effectiveness of the POC method is also found in the decoy sets from membrane protein loops. Furthermore, the POC method outperforms the other popularly-used consensus strategies in model ranking, such as rank-by-number, rank-by-rank, rank-by-vote, and regression-based methods. By integrating multiple knowledge- and physics-based scoring functions based on Pareto optimality and fuzzy dominance, the POC method is effective in distinguishing the best loop models from the other ones within a loop model set.
DOI: 10.1093/protein/gzn056
发表时间: 2008-12
期刊: Protein engineering, design & selection : PEDS
影响因子: --
作者:
Cui M;Mezei M;Osman R
通讯作者: Osman R
DOI: 10.1006/jmbi.1994.1109
发表时间: 1994-02-04
影响因子: 5.6
作者:
KOCHER, JPA;ROOMAN, MJ;WODAK, SJ
通讯作者: WODAK, SJ
DOI: 10.1021/ct800051k
发表时间: 2008-05-01
影响因子: 5.5
作者:
Felts, Anthony K.;Gallicchio, Emilio;Levy, Ronald M.
通讯作者: Levy, Ronald M.
DOI: 10.1002/prot.10613
发表时间: 2004-05-01
影响因子: 2.9
作者:
Jacobson, MP;Pincus, DL;Friesner, RA
通讯作者: Friesner, RA
DOI: 10.1186/1472-6807-9-28
发表时间: 2009-05-06
影响因子: --
作者:
Gao, Xin;Bu, Dongbo;Li, Ming
通讯作者: Li, Ming