PFRES: protein fold classification by using evolutionary information and predicted secondary structure

PFRES: protein fold classification by using evolutionary information and predicted secondary structure
复制标题

DOI:
10.1093/bioinformatics/btm475
复制
发表时间:
2007-11-01
期刊:
影响因子:
5.8
通讯作者:
Kurgan, Lukasz
Kurgan, Lukasz
中科院分区:
生物学3区
文献类型:
--
作者:
Chen, Ke;Kurgan, Lukasz

文献摘要

被引文献

相似文献

动机:蛋白质家族的数量估计只有1000个。最近的研究表明,沉积在PDB中的新结构的发现和SCOP类别的相关增长速度正在放缓。这表明蛋白质结构空间将很快被覆盖,因此我们可能能够通过使用已知的折叠模式推导出大多数剩余的结构。现有的三级结构预测方法在预测同源结构时效果良好,但在没有同源模板时预测效果较差。与此同时,一些具有模糊区序列特征的蛋白质可以形成类似的折叠。因此,确定结构相似性而不确定序列相似性有利于三级结构的预测。结果:本文提出的PFRES方法对低同源性(35个)序列进行蛋白质折叠自动分类,准确率分别为66.4%和68.4%。与现有方法相比,PFRES的准确率提高了6.3 ~ 12.4%。PFRES的预测精度显著优于竞争方法的预测精度。我们的方法采用了一个精心设计的、基于集成的分类器,以及一个新颖的、紧凑的、定制设计的特征表示,它比最准确的竞争方法(36比283)的特征表示少了近90%。该方法将基于PSI-BLAST剖面的演化信息与基于PSI-PRED预测的二级结构信息相结合。
Motivation: The number of protein families has been estimated to be as small as 1000. Recent study shows that the growth in discovery of novel structures that are deposited into PDB and the related rate of increase of SCOP categories are slowing down. This indicates that the protein structure space will be soon covered and thus we may be able to derive most of remaining structures by using the known folding patterns. Present tertiary structure prediction methods behave well when a homologous structure is predicted, but give poorer results when no homologous templates are available. At the same time, some proteins that share twilight-zone sequence identity can form similar folds. Therefore, determination of structural similarity without sequence similarity would be beneficial for prediction of tertiary structures.Results: The proposed PFRES method for automated protein fold classification from low identity (35) sequences obtains 66.4% and 68.4% accuracy for two test sets, respectively. PFRES obtains 6.3-12.4% higher accuracy than the existing methods. The prediction accuracy of PFRES is shown to be statistically significantly better than the accuracy of competing methods. Our method adopts a carefully designed, ensemble-based classifier, and a novel, compact and custom-designed feature representation that includes nearly 90% less features than the representation of the most accurate competing method (36 versus 283). The proposed representation combines evolutionary information by using the PSI-BLAST profile-based composition vector and information extracted from the secondary structure predicted with PSI-PRED.