Improving a consensus approach for protein structure selection by removing redundancy.

Improving a consensus approach for protein structure selection by removing redundancy.
复制标题

通过消除冗余改进蛋白质结构选择的共识方法。

DOI:
10.1109/tcbb.2011.75
复制
发表时间:
2011
期刊:
IEEE/ACM transactions on computational biology and bioinformatics
影响因子:
--
通讯作者:
Xu,Dong
Xu,Dong
中科院分区:
--
文献类型:
--
作者:
Wang,Qingguo;Shang,Yi;Xu,Dong

文献摘要

相似文献

在蛋白质三级结构预测中,关键的一步是从大量预测的结构模型中选择近天然结构。多年来,人们对蛋白质结构选择问题进行了广泛的研究,大多数方法都集中在开发更准确的能量或评分函数。尽管在这一领域取得了重大进展,但目前方法的识别能力仍然不能令人满意。在本文中,我们提出了一种新的基于共识的算法来选择预测的蛋白质结构。给定一组预测模型,我们的方法首先删除冗余结构,以获得参考模型的子集。然后,基于其与参考模型的平均成对相似性对结构进行排名。使用包含122个目标的大量预测模型的CASP 8数据集,我们将我们的方法与最好的CASP 8质量评估(QA)服务器进行了比较,这些服务器都是基于共识的,并表明我们的QA评分与GDT-TS的相关性优于CASP 8 QA服务器。我们还将我们的方法与最先进的评分函数进行了比较,并展示了其在近原生模型选择方面的改进性能。通过我们的方法选择的顶级模型的GDT-TS平均比最佳性能评分函数选择的模型高出8%以上。
In protein tertiary structure prediction, a crucial step is to select near-native structures from a large number of predicted structural models. Over the years, extensive research has been conducted for the protein structure selection problem with most approaches focusing on developing more accurate energy or scoring functions. Despite significant advances in this area, the discerning power of current approaches is still unsatisfactory. In this paper, we propose a novel consensus-based algorithm for the selection of predicted protein structures. Given a set of predicted models, our method first removes redundant structures to derive a subset of reference models. Then, a structure is ranked based on its average pairwise similarity to the reference models. Using the CASP8 data set containing a large collection of predicted models for 122 targets, we compared our method with the best CASP8 quality assessment (QA) servers, which are all consensus based, and showed that our QA scores correlate better with the GDT-TSs than those of the CASP8 QA servers. We also compared our method with the state-of-the-art scoring functions and showed its improved performance for near-native model selection. The GDT-TSs of the top models picked by our method are on average more than 8 percent better than the ones selected by the best performing scoring function.