PRED-CLASS: Cascading neural networks for generalized protein classification and genome-wide applications

PRED-CLASS: Cascading neural networks for generalized protein classification and genome-wide applications
复制标题

DOI:
10.1002/prot.1101
复制
发表时间:
2001-08-15
影响因子:
2.9
通讯作者:
Hamodrakas, SJ
Hamodrakas, SJ
中科院分区:
生物学4区
文献类型:
--
作者:
Pasquier, C;Promponas, VJ;Hamodrakas, SJ

文献摘要

被引文献

相似文献

一个级联系统的分层,人工神经网络(名为PRED-CLASS)的广义分类的蛋白质分为四个不同的类-跨膜,纤维,球状,和混合的信息单独编码在其氨基酸序列。单个组件网络的架构保持非常简单,减少了自由参数(网络突触权重)的数量,以实现更快的训练,提高泛化能力,并避免数据过拟合。从分布在四个目标类别(6个跨膜,10个纤维,13个球状和17个混合)中的少至50个蛋白质序列中捕获信息,PRED-CLASS能够从一组387个蛋白质中获得371个正确的预测(成功率接近96%),这些蛋白质明确地分配到一个目标类别中。PRED-CLASS的应用程序的几个测试集和完整的蛋白质组的几种生物体表明,这种方法可以作为一个有价值的工具,在注释基因组开放阅读框架,没有功能分配或作为一个初步的步骤,在折叠识别和从头结构预测方法。对于各种数据集和完整的基因组获得的详细结果,连同运行PRED-CLASS算法的网络服务器,沿着可以通过万维网在http://o2.biol.uoa.gr/PRED-CLASS上访问。(C)2001 Wiley-Liss,Inc.
A cascading system of hierarchical, artificial neural networks (named PRED-CLASS) is presented for the generalized classification of proteins into four distinct classes-transmembrane, fibrous, globular, and mixed-from information solely encoded in their amino acid sequences. The architecture of the individual component networks is kept very simple, reducing the number of free parameters (network synaptic weights) for faster training, improved generalization, and the avoidance of data overfitting. Capturing information from as few as 50 protein sequences spread among the four target classes (6 transmembrane, 10 fibrous, 13 globular, and 17 mixed), PRED-CLASS was able to obtain 371 correct predictions out of a set of 387 proteins (success rate similar to 96%) unambiguously assigned into one of the target classes. The application of PRED-CLASS to several test sets and complete proteomes of several organisms demonstrates that such a method could serve as a valuable tool in the annotation of genomic open reading frames with no functional assignment or as a preliminary step in fold recognition and ab initio structure prediction methods. Detailed results obtained for various data sets and completed genomes, along with a web sever running the PRED-CLASS algorithm, can be accessed over the World Wide Web at http://o2.biol.uoa.gr/PRED-CLASS. (C) 2001 Wiley-Liss, Inc.