RCPdb: An evolutionary classification and codon usage database for repeat-containing proteins

RCPdb: An evolutionary classification and codon usage database for repeat-containing proteins
复制标题

RCPdb: 含重复蛋白质的进化分类和密码子使用数据库

DOI:
10.1101/gr.6255407
复制
发表时间:
2007-07-01
期刊:
影响因子:
7
通讯作者:
Whisstock, James C.
Whisstock, James C.
中科院分区:
生物学1区
文献类型:
--
作者:
Faux, Noel G.;Huttley, Gavin A.;Whisstock, James C.

文献摘要

被引文献

相似文献

超过3%的人类蛋白质包含单一氨基酸重复序列(含重复序列的蛋白质,RCPs)。许多重复序列(同型肽)定位于参与转录的重要蛋白质中,并且某些重复序列的扩增,特别是多聚谷氨酰胺(poly - Q)和多聚丙氨酸(poly - A)序列,也可能导致神经系统疾病的发生。先前的研究表明,同型肽的组成是编码基因中富含G + C序列存在的结果,并且扩增是通过复制滑动发生的。在此,我们对13个物种中编码RCPs的基因变异进行了大规模的基因组分析,并将这些数据呈现在一个在线数据库(http://repeats.med.monash.edu.au/genetic_analysis/)中。这一资源允许对所考虑的真核生物物种中的RCPs、同型肽及其潜在的基因序列进行快速比较和分析。我们报告了三个主要发现。首先,在同型肽内重复的密码子只有一小部分存在偏向性,并且相对于生物体的转录组不存在G + C或A + T偏向性。其次,同型密码子的单碱基对颠换异常常见,这可能代表一种降低同型肽突变率的机制。第三,在不同物种间保守的同型肽位于受到更强纯化选择的区域,这与非保守的同型肽形成对比。
Over 3% of human proteins contain single amino acid repeats (repeat-containing proteins, RCPs). Many repeats (homopeptides) localize to important proteins involved in transcription, and the expansion of certain repeats, in particular poly-Q and poly-A tracts, can also lead to the development of neurological diseases. Previous studies have suggested that the homopeptide makeup is a result of the presence of G+C-rich tracts in the encoding genes and that expansion occurs via replication slippage. Here, we have performed a large-scale genomic analysis of the variation of the genes encoding RCPs in 13 species and present these data in an online database (http://repeats. med. monash. edu.au/genetic_analysis/). This resource allows rapid comparison and analysis of RCPs, homopeptides, and their underlying genetic tracts across the eukaryotic species considered. We report three major findings. First, there is a bias for a small subset of codons being reiterated within homopeptides, and there is no G+C or A+T bias relative to the organism's transcriptome. Second, single base pair transversions from the homocodon are unusually common and may represent a mechanism of reducing the rate of homopeptide mutations. Third, homopeptides that are conserved across different species lie within regions that are under stronger purifying selection in contrast to nonconserved homopeptides.