Evolutionary sequence modeling for discovery of peptide hormones.

Evolutionary sequence modeling for discovery of peptide hormones.
复制标题

DOI:
10.1371/journal.pcbi.1000258
复制
发表时间:
2009-01
影响因子:
4.3
通讯作者:
Toll L
Toll L
中科院分区:
生物学2区
文献类型:
--
作者:
Sonmez K;Zaveri NT;Kerman IA;Burke S;Neal CR;Xie X;Watson SJ;Toll L

文献摘要

参考文献

被引文献

相似文献

目前有大量的“孤儿”G蛋白偶联受体(GPCR),其内源性配体(肽激素)是未知的。这些肽类激素的鉴定是一个困难而重要的问题。我们描述了一个计算框架,模型的空间结构沿着的基因组序列,同时与跨物种的时间进化路径结构,并显示这样的模型可以用来发现新的功能分子,特别是肽激素,通过跨基因组序列比较。该计算框架将结构和进化约束的先验高级知识纳入进化概率模型的分层语法中。这种计算方法被用于确定新的激素原和加工肽的网站,通过在许多物种的功能元件水平上的序列比对。实验结果与算法的初始实现被用来识别潜在的激素原通过比较已知注释的蛋白质的Swiss-Prot数据库中的人类和非人类蛋白质。在这个概念证明中,我们在54种激素原中鉴定出45种,只有44种假阳性。已知的和假设的人类和小鼠蛋白质的比较,导致在识别一种新的推定的激素原与至少四个潜在的神经肽。最后,为了验证计算方法,我们提出了新的推定肽激素的基本分子生物学特征,包括其识别和区域定位在大脑中。这种基于HMM的物种比较计算方法成功地从全基因组蛋白质序列中识别出了一种以前未发现的神经肽。这种新的假定肽激素被发现在谨慎的大脑区域以及其他器官。这种方法的成功将对我们对GPCR和相关途径的理解产生重大影响,并有助于确定药物开发的新靶点。肽激素或神经肽由一串氨基酸组成,范围从大约3到50个残基。这些肽是由一种称为激素原的较大蛋白质加工而成,并激活一类称为G蛋白偶联受体(GPCR)的蛋白质。神经肽信号神经元和其他细胞导致细胞生物化学和潜在的基因表达的变化。存在许多“孤儿”GPCR,即,通过基因组测序或克隆发现的受体,其中其各自的肽激素是未知的。我们设计了一种计算方法,该方法可以同时模拟蛋白质序列中的模式和物种间的进化差异,以识别以前未知的肽激素。我们已经使用这种计算方法,以确定一个以前未知的假定激素原,其中包含多达四个潜在的神经肽,我们的特点是这种激素原在大鼠大脑和各种人体组织中的位置。这种计算技术将有助于识别其他神经肽,并有助于表征孤儿GPCR。由于大约一半的药物通过激活或抑制GPCR起作用,因此该技术应导致识别其他药物靶标并最终识别临床使用的药物。
There are currently a large number of “orphan” G-protein-coupled receptors (GPCRs) whose endogenous ligands (peptide hormones) are unknown. Identification of these peptide hormones is a difficult and important problem. We describe a computational framework that models spatial structure along the genomic sequence simultaneously with the temporal evolutionary path structure across species and show how such models can be used to discover new functional molecules, in particular peptide hormones, via cross-genomic sequence comparisons. The computational framework incorporates a priori high-level knowledge of structural and evolutionary constraints into a hierarchical grammar of evolutionary probabilistic models. This computational method was used for identifying novel prohormones and the processed peptide sites by producing sequence alignments across many species at the functional-element level. Experimental results with an initial implementation of the algorithm were used to identify potential prohormones by comparing the human and non-human proteins in the Swiss-Prot database of known annotated proteins. In this proof of concept, we identified 45 out of 54 prohormones with only 44 false positives. The comparison of known and hypothetical human and mouse proteins resulted in the identification of a novel putative prohormone with at least four potential neuropeptides. Finally, in order to validate the computational methodology, we present the basic molecular biological characterization of the novel putative peptide hormone, including its identification and regional localization in the brain. This species comparison, HMM-based computational approach succeeded in identifying a previously undiscovered neuropeptide from whole genome protein sequences. This novel putative peptide hormone is found in discreet brain regions as well as other organs. The success of this approach will have a great impact on our understanding of GPCRs and associated pathways and help to identify new targets for drug development. Peptide hormones, or neuropeptides, are made up of a string of amino acids ranging from approximately 3 to 50 residues. These peptides are processed from a larger protein called a prohormone and activate a class of proteins called G-protein-coupled receptors (GPCRs). Neuropeptides signal neurons and other cells leading to changes in cellular biochemistry and potentially gene expression. There are a number of “orphan” GPCRs, i.e., receptors that have been discovered either by genomic sequence or by cloning, in which its respective peptide hormone is unknown. We have devised a computational method that models patterns in protein sequence simultaneously with evolutionary differences across species in order to identify previously unknown peptide hormones. We have used this computational methodology to identify a previously unknown putative prohormone that contains up to four potential neuropeptides, and we have characterized this prohormone with respect to location in rat brain and various human tissues. This computational technique will be useful for the identification of additional neuropeptides and help to characterize orphan GPCRs. Because roughly half of all pharmaceuticals act through activation or inhibition of GPCRs, this technique should lead to the identification of additional pharmaceutical targets and ultimately clinically used drugs.
DOI: 10.1038/258577a0
发表时间: 1975-01-01
期刊: NATURE
影响因子: 64.8
作者:
HUGHES, J;SMITH, TW;MORRIS, HR
通讯作者: MORRIS, HR
DOI: 10.1006/jmbi.1994.1104
发表时间: 1994-02-04
影响因子: 5.6
作者:
KROGH, A;BROWN, M;HAUSSLER, D
通讯作者: HAUSSLER, D
DOI: 10.1038/nature01644
发表时间: 2003-05-15
期刊: NATURE
影响因子: 64.8
作者:
Kellis, M;Patterson, N;Lander, ES
通讯作者: Lander, ES
DOI: 10.1016/s0165-1838(99)00101-0
发表时间: 2000-03-15
期刊: JOURNAL OF THE AUTONOMIC NERVOUS SYSTEM
影响因子: --
作者:
Cano, G;Card, JP;Sved, AF
通讯作者: Sved, AF
DOI: 10.1093/bioinformatics/bth153
发表时间: 2004-08-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
McAuliffe, JD;Pachter, L;Jordan, MI
通讯作者: Jordan, MI