A NOVEL-APPROACH TO PREDICTING PROTEIN STRUCTURAL CLASSES IN A (20-1)-D AMINO-ACID-COMPOSITION SPACE

A NOVEL-APPROACH TO PREDICTING PROTEIN STRUCTURAL CLASSES IN A (20-1)-D AMINO-ACID-COMPOSITION SPACE
复制标题

DOI:
10.1002/prot.340210406
复制
发表时间:
1995-04-01
期刊:
PROTEINS-STRUCTURE FUNCTION AND GENETICS
影响因子:
--
通讯作者:
CHOU, KC
CHOU, KC
中科院分区:
其他
文献类型:
--
作者:
CHOU, KC

文献摘要

被引文献

相似文献

基于统计理论的预测方法的发展一般包括两部分:一是新算法的探索,二是训练数据库的完善。目前的研究致力于从两个方面改进蛋白质结构类别的预测,为了探索一种新的算法,开发了一种方法,通过协方差矩阵考虑蛋白质不同氨基酸成分之间的耦合效应,为了改进训练数据库,对蛋白质进行选择,使它们具有(1)尽可能多的非同源结构,并且 (2)结构质量好。因此,选择了129个代表性蛋白质。根据更好地反映相关结构类别特征的新标准,将它们分为30个α、30个β、30个α+β、30个α/β和9个zeta(不规则)蛋白质。当前方法对4×30个规则蛋白质的平均预测准确率为99.2%,对未包含在训练数据库中的64个独立测试蛋白质的平均预测准确率为95.3%。为了进一步验证其效率,对当前方法以及以前的方法进行了折刀分析,结果也非常有利于当前方法。为了完成数学基础,在附录A中提出并证明了一个定理,该定理对于更深层次地理解该新方法具有指导意义。 (C) 1995 Wiley-Liss, Inc.
The development of prediction methods based on statistical theory generally consists of two parts: one is focused on the exploration of new algorithms, and the other on the improvement of a training database. The current study is devoted to improving the prediction of protein structural classes from both of the two aspects, To explore a new algorithm, a method has been developed that makes allowance for taking into account the coupling effect among different amino acid components of a protein by a covariance matrix, To improve the training database, the selection of proteins is carried out so that they have (1) as many nonhomologous structures as possible, and (2) a good quality of structure. Thus, 129 representative proteins are selected. They are classified into 30 alpha, 30 beta, 30 alpha + beta, 30 alpha/beta, and 9 zeta (irregular) proteins according to a new criterion that better reflects the feature of the structural classes concerned, The average accuracy of prediction by the current method for the 4 x 30 regular proteins is 99.2%, and that for 64 independent testing proteins not included in the training database is 95.3%. To further validate its efficiency, a jackknife analysis has been performed for the current method as well as the previous ones, and the results are also much in favor of the current method, To complete the mathematical basis, a theorem is presented and proved in Appendix A that is instructive for understanding the novel method at a deeper level. (C) 1995 Wiley-Liss, Inc.