How well can the accuracy of comparative protein structure models be predicted?

How well can the accuracy of comparative protein structure models be predicted?
复制标题

DOI:
10.1110/ps.036061.108
复制
发表时间:
2008-11-01
期刊:
影响因子:
8
通讯作者:
Sali, Andrej
Sali, Andrej
中科院分区:
生物学3区
文献类型:
--
作者:
Eramian, David;Eswar, Narayanan;Sali, Andrej

文献摘要

被引文献

相似文献

比较结构模型可用于两个数量级以上的蛋白质序列比实验确定的结构。然而,这些模型有两个实验确定的结构所没有的限制:它们经常包含显著的误差,并且它们的准确性不能很容易地评估。我们已经解决了后者的限制,通过开发一个协议,专门用于预测的C α均方根偏差(RMSD)和本地重叠(NO3.5埃)的模型在其本地结构的情况下的错误进行了优化。与仅预测一个模型比其他模型更准确的大多数传统评估分数相比,这种方法在绝对意义上量化了误差,从而有助于确定模型是否适合预期应用。该评估依赖于由支持向量机构建的模型特定的评分函数。这种回归优化了多达九个特征的权重,包括各种序列相似性度量和统计潜力,这些特征是从被评估模型所特有的定制训练集中提取的:如果可能,我们使用具有相同倍数的类似大小的模型;否则,我们使用具有相同二级结构组成的类似大小的模型。该方案预测了6174个序列的580,317个比较模型的不同集合的RMSD和NO3.5埃误差,与实际误差的相关系数(r)分别为0.84和0.86。与其他13个测试的评估标准相比,该评分函数实现了最佳相关性,其相关性范围为0.35至0.71。
Comparative structure models are available for two orders of magnitude more protein sequences than are experimentally determined structures. These models, however, suffer from two limitations that experimentally determined structures do not: They frequently contain significant errors, and their accuracy cannot be readily assessed. We have addressed the latter limitation by developing a protocol optimized specifically for predicting the C alpha root-mean-squared deviation (RMSD) and native overlap (NO3.5 angstrom) errors of a model in the absence of its native structure. In contrast to most traditional assessment scores that merely predict one model is more accurate than others, this approach quantifies the error in an absolute sense, thus helping to determine whether or not the model is suitable for intended applications. The assessment relies on a model-specific scoring function constructed by a support vector machine. This regression optimizes the weights of up to nine features, including various sequence similarity measures and statistical potentials, extracted from a tailored training set of models unique to the model being assessed: If possible, we use similarly sized models with the same fold; otherwise, we use similarly sized models with the same secondary structure composition. This protocol predicts the RMSD and NO3.5 angstrom errors for a diverse set of 580,317 comparative models of 6174 sequences with correlation coefficients (r) of 0.84 and 0.86, respectively, to the actual errors. This scoring function achieves the best correlation compared to 13 other tested assessment criteria that achieved correlations ranging from 0.35 to 0.71.