Benchmarking consensus model quality assessment for protein fold recognition.

Benchmarking consensus model quality assessment for protein fold recognition.
复制标题

DOI:
10.1186/1471-2105-8-345
复制
发表时间:
2007-09-18
期刊:
影响因子:
3
通讯作者:
McGuffin LJ
McGuffin LJ
中科院分区:
生物学4区
文献类型:
--
作者:
McGuffin LJ

文献摘要

参考文献

被引文献

相似文献

从众多的备选方案中选择最高质量的蛋白质结构3D模型仍然是结构生物信息学领域的一个重要挑战。已经开发了许多模型质量评估程序(MQAP),其采用各种策略来解决这个问题,范围从能够基于单个模型产生单个能量分数的所谓“真正”MQAP到依赖于多个模型的结构比较或来自元服务器的附加信息的方法。然而,很明显,目前没有一种方法可以始终如一地将最高精度模型与最低精度模型分开。在本文中,一些表现最好的MQAP方法在它们添加到蛋白质折叠识别的潜在价值的背景下进行基准测试。还描述了两种新的方法:ModSSEA,它基于预测的二级结构元素的对齐和ModFOLD,它结合了几个真正的MQAP方法,使用人工神经网络。ModSSEA方法被认为是一种有效的模型质量评估程序,用于对来自许多服务器的多个模型进行排名,但是通过使用ModFOLD的共识方法可以获得更高的准确性。ModFOLD方法被证明显着优于真正的MQAP测试和竞争力的方法,利用集群或来自多个服务器的附加信息。几个真正的MQAP也被证明可以通过改进模型选择来为大多数单个折叠识别服务器增加价值,当作为后过滤器应用时,可以对模型进行重新排名。MQAP应该根据其预期使用的实际环境进行适当的基准测试。基于聚类的方法是性能最好的MQAP,其中许多模型可从许多服务器获得;然而,当有限的模型可用时,它们通常不会为单个折叠识别服务器增加价值。相反,测试的真正MQAP方法通常可以用作有效的后过滤器,用于对来自单个折叠识别服务器的少数模型进行重新排名,并且可以使用这些方法的共识来实现进一步的改进。
Selecting the highest quality 3D model of a protein structure from a number of alternatives remains an important challenge in the field of structural bioinformatics. Many Model Quality Assessment Programs (MQAPs) have been developed which adopt various strategies in order to tackle this problem, ranging from the so called "true" MQAPs capable of producing a single energy score based on a single model, to methods which rely on structural comparisons of multiple models or additional information from meta-servers. However, it is clear that no current method can separate the highest accuracy models from the lowest consistently. In this paper, a number of the top performing MQAP methods are benchmarked in the context of the potential value that they add to protein fold recognition. Two novel methods are also described: ModSSEA, which based on the alignment of predicted secondary structure elements and ModFOLD which combines several true MQAP methods using an artificial neural network. The ModSSEA method is found to be an effective model quality assessment program for ranking multiple models from many servers, however further accuracy can be gained by using the consensus approach of ModFOLD. The ModFOLD method is shown to significantly outperform the true MQAPs tested and is competitive with methods which make use of clustering or additional information from multiple servers. Several of the true MQAPs are also shown to add value to most individual fold recognition servers by improving model selection, when applied as a post filter in order to re-rank models. MQAPs should be benchmarked appropriately for the practical context in which they are intended to be used. Clustering based methods are the top performing MQAPs where many models are available from many servers; however, they often do not add value to individual fold recognition servers when limited models are available. Conversely, the true MQAP methods tested can often be used as effective post filters for re-ranking few models from individual fold recognition servers and further improvements can be achieved using a consensus of these methods.
DOI: 10.1016/s0022-2836(03)00323-1
发表时间: 2003-05-23
影响因子: 5.6
作者:
Keasar, C;Levitt, M
通讯作者: Levitt, M
DOI: 10.1093/bioinformatics/bti540
发表时间: 2005-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Pettitt, CS;McGuffin, LJ;Jones, DT
通讯作者: Jones, DT
DOI: 10.1110/ps.9.7.1399
发表时间: 2000-07-01
期刊: PROTEIN SCIENCE
影响因子: 8
作者:
Samudrala, R;Levitt, M
通讯作者: Levitt, M
DOI: 10.1006/jmbi.1996.0256
发表时间: 1996-05-03
影响因子: 5.6
作者:
Park, B;Levitt, M
通讯作者: Levitt, M
DOI: 10.1186/1471-2105-7-288
发表时间: 2006-06-07
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
McGuffin, Liam J.;Smith, Richard T.;Jones, David T.
通讯作者: Jones, David T.