SeqRate: sequence-based protein folding type classification and rates prediction.

SeqRate: sequence-based protein folding type classification and rates prediction.
复制标题

DOI:
10.1186/1471-2105-11-s3-s1
复制
发表时间:
2010-04-29
期刊:
影响因子:
3
通讯作者:
Cheng J
Cheng J
中科院分区:
生物学4区
文献类型:
--
作者:
Lin GN;Wang Z;Xu D;Cheng J

文献摘要

被引文献

相似文献

蛋白质折叠率是蛋白质的一个重要性质。预测蛋白质折叠速度有助于了解蛋白质折叠过程,指导蛋白质设计。大多数以前预测蛋白质折叠速度的方法都需要蛋白质的三级结构作为输入。而且大多数方法没有区分蛋白质的不同动力学性质(双态折叠或多态折叠)。在这里,我们开发了一种方法SeqRate,利用支持向量机从蛋白质序列预测的序列长度、氨基酸组成、接触顺序、接触数和二级结构信息来预测蛋白质折叠动力学类型(二态和多态)和真实折叠速度。系统地研究了个体特征对折叠率预测的贡献。在标准基准数据集上,折叠动力学类型分类的准确率为80%。在以10为底的对数尺度下,两态蛋白质文件夹的皮尔逊相关系数和预测折叠速率与实验折叠速率的平均绝对差值(SEC-1)分别为0.81和0.79,三态蛋白质文件夹的预测折叠速率和实验折叠速率的平均绝对差(SEC-1)分别为0.80和0.68。SeqRate是第一个基于序列的蛋白质折叠类型分类方法,它的折叠率预测精度比以往的基于序列的方法有所提高。它的性能可以通过其他信息进一步增强,例如基于结构的几何接触作为输入。Web服务器和预测折叠率的软件都可以在http://casp.rnet.missouri.edu/fold_rate/index.html.上公开获得
Protein folding rate is an important property of a protein. Predicting protein folding rate is useful for understanding protein folding process and guiding protein design. Most previous methods of predicting protein folding rate require the tertiary structure of a protein as an input. And most methods do not distinguish the different kinetic nature (two-state folding or multi-state folding) of the proteins. Here we developed a method, SeqRate, to predict both protein folding kinetic type (two-state versus multi-state) and real-value folding rate using sequence length, amino acid composition, contact order, contact number, and secondary structure information predicted from only protein sequence with support vector machines. We systematically studied the contributions of individual features to folding rate prediction. On a standard benchmark dataset, the accuracy of folding kinetic type classification is 80%. The Pearson correlation coefficient and the mean absolute difference between predicted and experimental folding rates (sec-1) in the base-10 logarithmic scale are 0.81 and 0.79 for two-state protein folders, and 0.80 and 0.68 for three-state protein folders. SeqRate is the first sequence-based method for protein folding type classification and its accuracy of fold rate prediction is improved over previous sequence-based methods. Its performance can be further enhanced with additional information, such as structure-based geometric contacts, as inputs. Both the web server and software of predicting folding rate are publicly available at http://casp.rnet.missouri.edu/fold_rate/index.html.