SCPRED: accurate prediction of protein structural class for sequences of twilight-zone similarity with predicting sequences.

SCPRED: accurate prediction of protein structural class for sequences of twilight-zone similarity with predicting sequences.
复制标题

DOI:
10.1186/1471-2105-9-226
复制
发表时间:
2008-05-01
期刊:
影响因子:
3
通讯作者:
Chen K
Chen K
中科院分区:
生物学4区
文献类型:
--
作者:
Kurgan L;Cios K;Chen K

文献摘要

参考文献

被引文献

相似文献

当预测同源蛋白质时,蛋白质结构预测方法可以提供准确的结果,而在没有同源模板的情况下获得较差的预测。然而,一些共享暮光区成对同一性的蛋白质链可以形成相似的折叠,因此在没有序列相似性的情况下确定结构相似性对于结构预测来说是理想的。蛋白质或其结构域的折叠类型被定义为结构类别。当前预测 SCOP 中定义的四个结构类别的结构类别预测方法为数据集提供高达 63% 的准确度,其中任何序列对的序列同一性属于暮光区。我们提出了 SCPRED 方法,该方法可以提高与用于预测的序列共享暮光区成对相似性的序列的预测准确性。 SCPRED 使用支持向量机分类器,该分类器采用多个定制设计的特征作为输入来预测结构类别。基于广泛的设计,考虑了超过 2300 个基于指数、成分和理化性质的特征以及基于预测二级结构和内容的特征,分类器的输入包括 8 个基于从 PSI-PRED 预测的二级结构中提取的信息的特征和一个根据序列计算的特征。使用 1673 条蛋白质链的数据集(其中任何一对序列都具有暮光区相似性)进行的测试表明,SCPRED 在预测 SCOP 定义的四个结构类别时获得了 80.3% 的准确率,这比最近十几种基于支持向量机、逻辑回归和分类器预测器集成的竞争方法要优越。 SCPRED 可以准确地找到与用于预测的序列具有较低同一性的序列的相似结构。 SCPRED 实现的高预测精度归功于特征的设计,尽管其维度较低,但仍能够分离结构类别。我们还证明了 SCPRED 的预测可以成功地用作后处理过滤器,以提高现代折叠分类方法的性能。
Protein structure prediction methods provide accurate results when a homologous protein is predicted, while poorer predictions are obtained in the absence of homologous templates. However, some protein chains that share twilight-zone pairwise identity can form similar folds and thus determining structural similarity without the sequence similarity would be desirable for the structure prediction. The folding type of a protein or its domain is defined as the structural class. Current structural class prediction methods that predict the four structural classes defined in SCOP provide up to 63% accuracy for the datasets in which sequence identity of any pair of sequences belongs to the twilight-zone. We propose SCPRED method that improves prediction accuracy for sequences that share twilight-zone pairwise similarity with sequences used for the prediction. SCPRED uses a support vector machine classifier that takes several custom-designed features as its input to predict the structural classes. Based on extensive design that considers over 2300 index-, composition- and physicochemical properties-based features along with features based on the predicted secondary structure and content, the classifier's input includes 8 features based on information extracted from the secondary structure predicted with PSI-PRED and one feature computed from the sequence. Tests performed with datasets of 1673 protein chains, in which any pair of sequences shares twilight-zone similarity, show that SCPRED obtains 80.3% accuracy when predicting the four SCOP-defined structural classes, which is superior when compared with over a dozen recent competing methods that are based on support vector machine, logistic regression, and ensemble of classifiers predictors. The SCPRED can accurately find similar structures for sequences that share low identity with sequence used for the prediction. The high predictive accuracy achieved by SCPRED is attributed to the design of the features, which are capable of separating the structural classes in spite of their low dimensionality. We also demonstrate that the SCPRED's predictions can be successfully used as a post-processing filter to improve performance of modern fold classification methods.
DOI: 10.1093/bioinformatics/btm475
发表时间: 2007-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Chen, Ke;Kurgan, Lukasz
通讯作者: Kurgan, Lukasz
DOI: 10.1016/j.jtbi.2005.05.034
发表时间: 2006-01-07
影响因子: 2
作者:
Cai, YD;Feng, KY;Chou, KC
通讯作者: Chou, KC
DOI: 10.1093/bioinformatics/btl453
发表时间: 2006-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Birzele, Fabian;Kramer, Stefan
通讯作者: Kramer, Stefan
DOI: 10.1093/nar/gkh039
发表时间: 2004-01-01
影响因子: 14.9
作者:
Andreeva, A;Howorth, D;Murzin, AG
通讯作者: Murzin, AG
DOI: 10.1186/1471-2105-2-3
发表时间: 2001
期刊: BMC bioinformatics
影响因子: 3
作者:
Cai YD;Liu XJ;Xu X;Zhou GP
通讯作者: Zhou GP