The protein structure prediction problem could be solved using the current PDB library

The protein structure prediction problem could be solved using the current PDB library
复制标题

DOI:
10.1073/pnas.0407152101
复制
发表时间:
2005-01-25
影响因子:
11.1
通讯作者:
Skolnick, J
Skolnick, J
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Zhang, Y;Skolnick, J

文献摘要

被引文献

相似文献

对于单域蛋白质,我们检查当前蛋白质数据库(PDB)库中结构的完整性,以用于未知序列的全长模型构建。为了解决这个问题,我们采用了包含 1,489 个中等大小蛋白质的综合基准集,以 35% 的序列同一性水平覆盖 PDB,并通过结构比对来识别模板。排除同源蛋白后,我们总能找到与天然蛋白相似的折叠,与天然蛋白的平均均方根偏差 (RMSD) 为 2.5 埃,比对覆盖率约为 82%。这些模板结构通常包含大量插入/删除。 TASSER 算法用于构建全长模型,其中连续片段从得分最高的模板中切除,并在优化力场的指导下重新组装,其中包括从模板中获取的共识约束和基于知识的统计势。对于几乎所有目标(2/1,489 除外),生成的全长模型的 RMSD 低于 6 埃(其中 97% 低于 4 埃)。平均而言,全长模型的 RMSD 为 2.25 埃,对齐区域从 2.5 A 提高到 1.88 埃,与低分辨率实验结构的准确性相当。此外,从最先进的结构比对出发,我们演示了一种能够始终如一地使基于模板的比对更接近原生的方法。这些结果强烈表明,原则上可以基于当前的 PDB 库,通过开发可以恢复此类初始对齐的有效折叠识别算法来解决蛋白质折叠问题。
For single-domain proteins, we examine the completeness of the structures in the current Protein Data Bank (PDB) library for use in full-length model construction of unknown sequences. To address this issue, we employ a comprehensive benchmark set of 1,489 medium-size proteins that cover the PDB at the level of 35% sequence identity and identify templates by structure alignment. With homologous proteins excluded, we can always find similar folds to native with an average rms deviation (RMSD) from native of 2.5 Angstrom with approximate to 82% alignment coverage. These template structures often contain a significant number of insertions/deletions. The TASSER algorithm was applied to build full-length models, where continuous fragments are excised from the top-scoring templates and reassembled under the guide of an optimized force field, which includes consensus restraints taken from the templates and knowledge-based statistical potentials. For almost all targets (except for 2/1,489), the resultant full-length models have an RMSD to native below 6 Angstrom (97% of them below 4 Angstrom). On average, the RMSD of full-length models is 2.25 Angstrom, with aligned regions improved from 2.5 A to 1.88 Angstrom, comparable with the accuracy of low-resolution experimental structures. Furthermore, starting from state-of-the-art structural alignments, we demonstrate a methodology that can consistently bring template-based alignments closer to native. These results are highly suggestive that the protein-folding problem can in principle be solved based on the current PDB library by developing efficient fold recognition algorithms that can recover such initial alignments.