Molecular fossils in the human genome: Identification and analysis of the pseudogenes in chromosomes 21 and 22

Molecular fossils in the human genome: Identification and analysis of the pseudogenes in chromosomes 21 and 22
复制标题

DOI:
10.1101/gr.207102
复制
发表时间:
2002-02-01
期刊:
影响因子:
7
通讯作者:
Gerstein, M
Gerstein, M
中科院分区:
生物学1区
文献类型:
--
作者:
Harrison, PM;Hegyi, H;Gerstein, M

文献摘要

被引文献

相似文献

我们已经开发了一种初步的方法来注释和调查人类基因组中的假基因。我们搜索人类基因组DNA,寻找与已知蛋白质序列相似且包含明显残缺的区域(即中序停止密码子或移码),同时确保与已知基因的注释尽量减少重叠。假基因可分为“已加工”和“未加工”;前者是从信使核糖核酸反转录而来(因此没有内含子结构),而后者可能是由基因组复制而来。我们根据是否有一个连续的同源性跨度,即最接近匹配的人类蛋白质长度的70%(即,去除内含子),或者是否有证据表明有多聚腺苷作用,来注释假定的已加工假基因。我们将我们的方法应用于21号和22号染色体,这是人类基因组的第一部分,完全测序,在测序中心报告的264个假基因注释之外,发现了190个新的假基因注释。在第21和22号染色体上,共有189个已加工假基因、195个未加工假基因,另外还有70个假基因片段。(详细!作业可在http://bioinfo.mbb.yale.edu/genome/pseudogene或http://genecensus.org/Posogene上查阅。)通过外推,我们预测在整个人类基因组中可能有多达20,000个类似的伪基因,其中略多于一半的假基因正在处理中。我们已经确定了21号和22号染色体上的主要群体和假基因簇。假基因相对于两条染色体着丝粒附近的基因有明显的过剩,表明基因组中存在假基因“热点”。我们已经研究了InterPro家族和基因本体论(GO)功能类别在我们的伪基因中的分布。总体而言,加工和未加工假基因群体中的家系都符合与基因家族相似的幂规律分布,有几个大家族和许多小家族。尤其是,经过处理的群体富含高表达的核糖体蛋白序列(类似于20%),这些序列似乎相当均匀地分布在染色体上。我们比较了不同进化年代的加工假基因,观察到“古代”和“现代”亚群之间高度相似。这可能归因于核糖体蛋白在进化过程中的持续高表达。最后,我们发现22号染色体的假基因群体是由免疫球蛋白片段主导的,与其他假基因群体相比,免疫球蛋白片段的每个氨基酸的致残率更高,而且分化程度也更大。
We have developed an initial approach for annotating and surveying pseudogenes in the human genome. We search human genomic DNA for regions that are similar to known protein sequences and contain obvious disablements (i.e., mid-sequence stop codons or frameshifts), while ensuring minimal overlap with annotations of known genes. Pseudogenes can be divided into "processed" and "non processed"; the former are reverse transcribed from mRNA (and therefore have no intron structure), whereas the latter presumably arise from genomic duplications. We annotate putative processed pseudogenes based on whether there is a continuous span of homology that is >70% of the length of the closest matching human protein (i.e., with introns removed), or whether there is evidence of polyadenylation. We have applied our approach to chromosomes 21 and 22, the first parts of the human genome completely sequenced, finding 190 new pseudogene annotations beyond the 264 reported by the sequencing centers. In total, on chromosomes 21 and 22, there are 189 processed pseudogenes, 195 nonprocessed pseudogenes, and, additionally, 70 pseudogenic immunoglobulin gene segments. (Detailed! assignments are available at http://bioinfo.mbb.yale.edu/genome/pseudogene or http:// genecensus.org/pseudogene.) By extrapolation, we predict that there could be up to similar to20,000 pseudogenes in the whole human genome, with a little more than half of them processed. We have determined the main populations and clusters of pseudogenes on chromosomes 21 and 22. There are notable excesses of pseudogenes relative to genes near the centromeres of both chromosomes, indicating the existence of pseudogenic "hot-spots" in the genome. We have looked at the distribution of InterPro families and Gene Ontology (GO) functional categories in our pseudogenes. Overall, the families in both processed and nonprocessed pseudogene populations occur according to a similar power-law distribution as that found for the occurrence of gene families, with a few big families and many small ones. The processed population is, in particular, enriched in highly expressed ribosomal-protein sequences (similar to20%), which appear fairly evenly distributed across the chromosomes. We compared processed pseudogenes of different evolutionary ages, observing a high degree of similarity between "ancient" and "modern" subpopulations. This may be attributable to the consistently high expression of ribosomal proteins over evolutionary time. Finally, we find that chromosome 22 pseudogene population is dominated by immunoglobulin segments, which have a greater rate of disablement per amino acid than the other pseudogene populations and are also substantially more diverged.