Estimating quality of template-based protein models by alignment stability

Estimating quality of template-based protein models by alignment stability
复制标题

DOI:
10.1002/prot.21819
复制
发表时间:
2008-05-15
影响因子:
2.9
通讯作者:
Kihara, Daisuke
Kihara, Daisuke
中科院分区:
生物学4区
文献类型:
--
作者:
Chen, Hao;Kihara, Daisuke

文献摘要

被引文献

相似文献

蛋白质三级结构预测中的误差是不可避免的,但目前的大多数预测算法都没有明确地显示出误差。预测结构的估计误差是实验生物学家使用预测模型设计和解释实验的关键信息。在这里,我们提出了一种方法来估计预测结构的最佳目标模板比对的稳定性的基础上,当与一组次优的比对。最佳比对的稳定性通过称为亚最佳比对多样性(SPAD)的指数来量化。我们实现了SPAD在一个基于配置文件的线程算法,并研究如何以及SPAD可以指示线程模型中的错误,使用一个大型的基准数据集的5232对齐。SPAD显示出非常好的相关性,不仅对齐移位误差,但也结构水平的错误,预测的结构模型的均方根偏差(RMSD)的天然结构(即全局误差),并在每个残基位置的局部误差。我们进一步比较了SPAD与其他七个质量指标,六个从基于序列的测量和一个原子的统计潜力,离散优化的蛋白质能量(DOPE),在相关系数的全局和局部结构水平的错误。在与结构模型RMSD的相关性方面,当靶和模板在同一个SCOP家族中时,序列同一性与RMSD的相关性最好,在超家族水平上,SPAD的相关性最好;在折叠水平上,DOPE的相关性最好。然而,在一个头对头的比较,SPAD赢了其他措施。接下来,SPAD与其他三种局部误差测量进行比较。在这一比较中,SPAD在所有的家庭,超家庭和折叠水平上都是最好的。利用发现的相关性,我们还预测了我们预测的结构的CASP 7目标的SPAD的全局和局部误差。最后,我们提出了一个香肠表示的预测三级结构,直观地显示预测的结构和估计的误差范围的结构。
The error in protein tertiary structure prediction is unavoidable, but it is not explicitly shown in most of the current prediction algorithms. Estimated error of a predicted structure is crucial information for experimental biologists to use the prediction model for design and interpretation of experiments. Here, we propose a method to estimate errors in predicted structures based on the stability of the optimal target-template alignment when compared with a set of suboptimal alignments. The stability of the optimal alignment is quantified by an index named the SuboPtimal Alignment Diversity (SPAD). We implemented SPAD in a profile-based threading algorithm and investigated how well SPAD can indicate errors in threading models using a large benchmark dataset of 5232 alignments. SPAD shows a very good correlation not only to alignment shift errors but also structure-level errors, the root mean square deviation (RMSD) of predicted structure models to the native structures (i.e. global errors), and local errors at each residue position. We have further compared SPAD with seven other quality measures, six from sequence alignment-based measures and one atomic statistical potential, discrete optimized protein energy (DOPE), in terms of the correlation coefficient to the global and local structure-level errors. In terms of the correlation to the RMSD of structure models, when a target and a template are in the same SCOP family, the sequence identity showed a best correlation to the RMSD, in the superfamily level, SPAD was the best; and in the fold level, DOPE was best. However, in a head-to-head comparison, SPAD wins over the other measures. Next, SPAD is compared with three other measures of local errors. In this comparison, SPAD was best in all of the family, the superfamily and the fold levels. Using the discovered correlation, we have also predicted the global and local error of our predicted structures of CASP7 targets by the SPAD. Finally, we proposed a sausage representation of predicted tertiary structures which intuitively indicate the predicted structure and the estimated error range of the structure simultaneously.