Identification and analysis of over 2000 ribosomal protein pseudogenes in the human genome

Identification and analysis of over 2000 ribosomal protein pseudogenes in the human genome
复制标题

DOI:
10.1101/gr.331902
复制
发表时间:
2002-10-01
期刊:
影响因子:
7
通讯作者:
Gerstein, M
Gerstein, M
中科院分区:
生物学1区
文献类型:
--
作者:
Zhang, ZL;Harrison, P;Gerstein, M

文献摘要

被引文献

相似文献

哺乳动物有79种核糖体蛋白(RP)。利用基于序列同源性的系统程序,我们在人类基因组中全面鉴定了这些蛋白质的假基因。我们的作业可以在http://www.pseudogene.org或http://bioinfo.mbb.yale.edu/genome/pseudogene上找到。共发现2090个加工过的假基因和16个RP基因的重复。相对于匹配的亲本蛋白,每个加工伪基因的平均相对序列长度为97%,平均序列同源性为76%。其中一小部分(258个)不包含明显的失能(停止密码子或帧移),因此可能被误认为是功能基因,178个被一个或多个重复元件破坏。平均而言,加工过的假基因在S‘端比3’端截断更长,这与靶引物-反转录(target-引物-reverse-transcription, TPRT)机制一致。有趣的是,在第16号染色体上,在一个功能性RPS2基因的内含子区域发现了一个RPL26加工的假基因。RP假基因在整个基因组中的大规模分布似乎主要是由于每条染色体上随机插入的数量,因此,与它的大小成正比。与RP基因相比,RP假基因在基因组gc -中间区域的密度最高(41%-46%),密度模式介于LINEs和Alus之间。这可以用负选择理论来解释,因为我们观察到富含gc的RP假基因在gc贫乏的区域衰减得更快。此外,我们观察到加工的假基因数量与相关功能基因的GC含量之间存在相关性,即相对GC-poor的rp具有更多加工的假基因。从RPL21的145个假基因到RPL14的3个假基因不等。根据RP假基因与现今RP基因的序列差异,我们能够确定RP假基因的年代,发现其年龄分布与Alus相似。这种分布与近40万年来原始人谱系中逆转录活动的下降是一致的。基于这些新发现,我们讨论了反转录转座子稳定性和基因组动力学的意义。
Mammals have 79 ribosomal proteins (RP). Using a systematic procedure based on sequence-homology, we have comprehensively identified pseudogenes of these proteins in the human genome. Our assignments are available at http://www.pseudogene.org or http://bioinfo.mbb.yale.edu/genome/pseudogene. In total, we found 2090 processed pseudogenes and 16 duplications of RP genes. In relation to the matching parent protein, each of the processed pseudogenes has an average relative sequence length of 97% and an average sequence identity of 76%. A small number (258) of them do not contain obvious disablements (stop codons or frameshifts) and, therefore, could be mistaken as functional genes, and 178 are disrupted by one or more repetitive elements. On average, processed pseudogenes have a longer truncation at the S' end than the 3' end, consistent with the target-primed-reverse-transcription (TPRT) mechanism. Interestingly, on chromosome 16, an RPL26 processed pseudogene was found in the intron region of a functional RPS2 gene. The large-scale distribution of RP pseudogenes throughout the genome appears to result, chiefly, from random insertions with the numbers on each chromosome, consequently, proportional to its size. In contrast to RP genes, the RP pseudogenes have the highest density in GC-intermediate regions (41%-46%) of the genome, with the density pattern being between that of LINEs and Alus. This can be explained by a negative selection theory as we observed that GC-rich RP pseudogenes decay faster in GC-poor regions. Also, we observed a correlation between the number of processed pseudogenes and the GC content of the associated functional gene, i.e., relatively GC-poor RPs have more processed pseudogenes. This ranges from 145 pseudogenes for RPL21 down to 3 pseudogenes for RPL14. We were able to date the RP pseudogenes based on their sequence divergence from present-day RP genes, finding an age distribution similar to that for Alus. The distribution is consistent with a decline in retrotransposition activity in the hominid lineage during the last 40 Myr. We discuss the implications for retrotransposon stability and genome dynamics based on these new findings.