Predicting functional family of novel enzymes irrespective of sequence similarity: a statistical learning approach.

Predicting functional family of novel enzymes irrespective of sequence similarity: a statistical learning approach.
复制标题

DOI:
10.1093/nar/gkh984
复制
发表时间:
2004
影响因子:
14.9
通讯作者:
Chen YZ
Chen YZ
中科院分区:
生物学2区
文献类型:
--
作者:
Han LY;Cai CZ;Ji ZL;Cao ZW;Cui J;Chen YZ

文献摘要

参考文献

被引文献

相似文献

没有已知功能的序列同源的蛋白质的功能很难根据序列相似性来分配。如果一个是新发现的,而另一个是唯一已知的相似序列的蛋白质,那么不同功能的同源蛋白质可能会出现同样的问题。探索不基于序列相似性的方法是可取的。一种方法是指定蛋白质的功能家族,以提供关于其功能的有用提示。几个小组已经采用了一种统计学习方法,支持向量机(SVMs),用于直接从序列预测蛋白质功能家族,而不考虑序列的相似性。这些研究表明,支持向量机的预测精度达到了对功能族分配有用的水平。但它对远亲蛋白质和不同功能的同源蛋白质的指配能力还没有得到严格和充分的评估。在这里,支持向量机被用于两组酶的功能家族分配。其中一种由50种酶组成,这些酶在蛋白质数据库的PSI-BLAST搜索中没有已知功能的同源物。另一组含有8对不同家族的同源酶。支持向量机正确地分配了第一组中72%的酶和第二组中62%的酶对,这表明它对于促进新蛋白质的功能研究是潜在的有用的。我们的软件SVMProt的网页版本可在http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi.上访问
The function of a protein that has no sequence homolog of known function is difficult to assign on the basis of sequence similarity. The same problem may arise for homologous proteins of different functions if one is newly discovered and the other is the only known protein of similar sequence. It is desirable to explore methods that are not based on sequence similarity. One approach is to assign functional family of a protein to provide useful hint about its function. Several groups have employed a statistical learning method, support vector machines (SVMs), for predicting protein functional family directly from sequence irrespective of sequence similarity. These studies showed that SVM prediction accuracy is at a level useful for functional family assignment. But its capability for assignment of distantly related proteins and homologous proteins of different functions has not been critically and adequately assessed. Here SVM is tested for functional family assignment of two groups of enzymes. One consists of 50 enzymes that have no homolog of known function from PSI-BLAST search of protein databases. The other contains eight pairs of homologous enzymes of different families. SVM correctly assigns 72% of the enzymes in the first group and 62% of the enzyme pairs in the second group, suggesting that it is potentially useful for facilitating functional study of novel proteins. A web version of our software, SVMProt, is accessible at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi.
DOI: 10.1128/jvi.78.4.2114-2120.2004
发表时间: 2004-02-01
影响因子: 5.4
作者:
Makeyev, EV;Bamford, DH
通讯作者: Bamford, DH
DOI: 10.1023/a:1009715923555
发表时间: 1998-06-01
影响因子: 4.8
作者:
Burges, CJC
通讯作者: Burges, CJC
DOI: 10.1016/s0022-2836(02)00379-0
发表时间: 2002-06-21
影响因子: 5.6
作者:
Jensen, LJ;Gupta, R;Brunak, S
通讯作者: Brunak, S
DOI: 10.1006/jmbi.2001.4580
发表时间: 2001-04-27
影响因子: 5.6
作者:
Hua, SJ;Sun, ZR
通讯作者: Sun, ZR
DOI: 10.1261/rna.5890304
发表时间: 2004-03-01
期刊: RNA
影响因子: 4.5
作者:
Han, LY;Cai, CZ;Chen, YZ
通讯作者: Chen, YZ