GPCR-GIA: a web-server for identifying G-protein coupled receptors and their families with grey incidence analysis

GPCR-GIA: a web-server for identifying G-protein coupled receptors and their families with grey incidence analysis
复制标题

GPCR-GIA:通过灰色关联分析识别 G 蛋白偶联受体及其家族的网络服务器

DOI:
10.1093/protein/gzp057
复制
发表时间:
2009-11-01
影响因子:
2.4
通讯作者:
Chou, Kuo-Chen
Chou, Kuo-Chen
中科院分区:
生物学4区
文献类型:
--
作者:
Lin, Wei-Zhong;Xiao, Xuan;Chou, Kuo-Chen

文献摘要

被引文献

相似文献

G蛋白偶联受体(GPCR)在调节各种生理过程以及几乎所有细胞的活性中发挥重要作用。不同的GPCR家族负责不同的功能。随着后基因组时代产生的蛋白质序列的雪崩,人们非常希望开发一种自动化方法来解决两个问题:给定查询蛋白质的序列,我们能否识别它是否是GPCR?如果是的话,它属于哪个家族?在这里,一个两层集成分类器称为GPCR-GIA提出了一种新的尺度称为“灰色关联度”。GPCR-GIA对GPCR和non-GPCR的识别成功率约为95%,对9个家系中GPCR的识别成功率约为80%。这些比率是通过在严格的基准数据集上的刀切交叉验证测试获得的,其中没有一种蛋白质与同一类中的任何其他蛋白质具有>= 50%的成对序列同一性。此外,还在http://218.65.61.89:8080/bioinfo/GPCR-GIA上建立了一个方便用户的网络服务器。为方便用户使用,提供了如何使用GPCR-GIA Web服务器的分步指南。一般来说,对于300-400个氨基酸的查询蛋白质序列,人们可以在大约10秒内得到期望的两级结果;序列越长,所需的时间越多。
G-protein-coupled receptors (GPCRs) play fundamental roles in regulating various physiological processes as well as the activity of virtually all cells. Different GPCR families are responsible for different functions. With the avalanche of protein sequences generated in the post-genomic age, it is highly desired to develop an automated method to address the two problems: given the sequence of a query protein, can we identify whether it is a GPCR? If it is, what family class does it belong to? Here, a two-layer ensemble classifier called GPCR-GIA was proposed by introducing a novel scale called 'grey incident degree'. The overall success rate by GPCR-GIA in identifying GPCR and non-GPCR was about 95%, and that in identifying the GPCRs among their nine family classes was about 80%. These rates were obtained by the jackknife cross-validation tests on the stringent benchmark data sets where none of the proteins has >= 50% pairwise sequence identity to any other in a same class. Moreover, a user-friendly web-server was established at http://218.65.61.89:8080/bioinfo/GPCR-GIA. For user's convenience, a step-by-step guide on how to use the GPCR-GIA web server is provided. Generally speaking, one can get the desired two-level results in around 10 s for a query protein sequence of 300-400 amino acids; the longer the sequence is, the more time that is needed.