On the hierarchical classification of G protein-coupled receptors

On the hierarchical classification of G protein-coupled receptors
复制标题

DOI:
10.1093/bioinformatics/btm506
复制
发表时间:
2007-12-01
期刊:
影响因子:
5.8
通讯作者:
Flower, Darren R.
Flower, Darren R.
中科院分区:
生物学3区
文献类型:
--
作者:
Davies, Matthew N.;Secker, Andrew;Flower, Darren R.

文献摘要

被引文献

相似文献

动机:G蛋白偶联受体(GPCRs)通过将细胞外信号转导为细胞内反应在许多生理系统中发挥重要作用。在所有上市的药物中,有50多种是针对GPCR的。有相当大的兴趣在开发一种算法,可以有效地预测从其主要序列的GPCR的功能。这样的算法是有用的,不仅在确定新的GPCR序列,但在表征已知的GPCRs.Results之间的相互关系:一个无障碍的方法来GPCR分类已开发使用的技术从数据挖掘和蛋白化学计量学。构建了超过8000个序列的数据集来训练算法。这是目前最大的气相化学还原数据集之一。基于蛋白质理化性质的最简单合理的数值表示,开发了一种预测算法。开发了一种自上而下的选择性方法,该方法使用分级分类器将序列分配给气相化学还原分级结构内的细分部分。该算法的预测性能进行了评估,对几个标准的数据挖掘分类器,并进一步验证支持向量机为基础的GPCR预测服务器。选择性自上而下的方法在几乎所有情况下都比标准数据挖掘方法实现了更高的准确性。
Motivation: G protein-coupled receptors (GPCRs) play an important role in many physiological systems by transducing an extracellular signal into an intracellular response. Over 50 of all marketed drugs are targeted towards a GPCR. There is considerable interest in developing an algorithm that could effectively predict the function of a GPCR from its primary sequence. Such an algorithm is useful not only in identifying novel GPCR sequences but in characterizing the interrelationships between known GPCRs.Results: An alignment-free approach to GPCR classification has been developed using techniques drawn from data mining and proteochemometrics. A dataset of over 8000 sequences was constructed to train the algorithm. This represents one of the largest GPCR datasets currently available. A predictive algorithm was developed based upon the simplest reasonable numerical representation of the proteins physicochemical properties. A selective top-down approach was developed, which used a hierarchical classifier to assign sequences to subdivisions within the GPCR hierarchy. The predictive performance of the algorithm was assessed against several standard data mining classifiers and further validated against Support Vector Machine-based GPCR prediction servers. The selective top-down approach achieves significantly higher accuracy than standard data mining methods in almost all cases.