Automated structure prediction of weakly homologous proteins on a genomic scale'

Automated structure prediction of weakly homologous proteins on a genomic scale'
复制标题

DOI:
10.1073/pnas.0305695101
复制
发表时间:
2004-05-18
影响因子:
11.1
通讯作者:
Skolnick, J
Skolnick, J
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Zhang, Y;Skolnick, J

文献摘要

被引文献

相似文献

我们开发了 TASSER,一种蛋白质结构预测的分层方法,包括通过线程识别模板,然后通过优化的 C 引导的连续模板片段重排进行三级结构组装,以及由基于线程的预测三级限制驱动的基于侧链的势。 TASSER 应用于蛋白质数据库中包含 1,489 种中等大小蛋白质的综合基准集。排除同源物后,在 927 个案例中,我们的线程算法 PROSPECTOR_3 识别的模板与原始模板的均方根偏差 < 6.5 埃,比对覆盖率接近 80%。模板重新组装后,这个数字增加到 1,172。这表明最终模型相对于初始模板对齐有显着和系统的改进。此外,还证明了循环建模的显着改进。然后,我们将 TASSER 应用于大肠杆菌基因组中的 1,360 个中等大小的 ORF;根据蛋白质数据库基准中建立的置信标准,可以高精度预测接近 920 的值。这些来自我们对所有蛋白质类别的前所未有的综合折叠基准的结果为 TASSER 在结构基因组学中的应用提供了可靠的基础,特别是对于低序列同一性的蛋白质以解决蛋白质结构。
We have developed TASSER, a hierarchical approach to protein structure prediction that consists of template identification by threading, followed by tertiary structure assembly via the rearrangement of continuous template fragments guided by an optimized C. and side-chain-based potential driven by threading based, predicted tertiary restraints. TASSER was applied to a comprehensive benchmark set of 1,489 medium-sized proteins in the Protein Data Bank. With homologues excluded, in 927 cases, the templates identified by our threading algorithm PROSPECTOR_3 have a rms deviation from native < 6.5 Angstrom with approximate to 80% alignment coverage. After template reassembly, this number increases to 1,172. This shows significant and systematic improvement of the final models with respect to the initial template alignments. Furthermore, significant improvements in loop modeling are demonstrated. We then apply TASSER to the 1,360 medium-sized ORFs in the Escherichia coli genome; approximate to 920 can be predicted with high accuracy based on confidence criteria established in the Protein Data Bank benchmark. These results from our unprecedented comprehensive folding benchmark on all protein categories provide a reliable basis for the application Of TASSER to structural genomics, especially to proteins of low sequence identity to solved protein structures.