Reevaluating human gene annotation: A second-generation analysis of chromosome 22

Reevaluating human gene annotation: A second-generation analysis of chromosome 22
复制标题

DOI:
10.1101/gr.695703
复制
发表时间:
2003-01-01
期刊:
影响因子:
7
通讯作者:
Dunham, I
Dunham, I
中科院分区:
生物学1区
文献类型:
--
作者:
Collins, JE;Goward, ME;Dunham, I

文献摘要

被引文献

相似文献

我们报道了人类22号染色体的第二代基因注释。利用表达序列数据库,比较序列分析和实验验证,我们扩展了基因,融合了以前的片段结构,并鉴定了新的基因。该注释的外显子总长度比我们之前发表的注释增加了74%,包含546个蛋白质编码基因和234个假基因。32个潜在的蛋白质编码注释是其他基因的部分拷贝,可能代表了所有进化路径上的复制,以改变或丧失功能。我们还鉴定了31种非蛋白质编码转录本,包括16种可能的反义rna。通过外推,我们估计人类基因组包含29,000-36,000个蛋白质编码基因,21,300个假基因和1,500个反义rna。我们认为,我们修订的注释标准为未来人类基因组的注释提供了一个范例。
We report a second-generation gene annotation of human chromosome 22. Using expressed sequence databases, comparative sequence analysis, and experimental verification, we have extended genes, fused previously fragmented structures, and identified new genes. The total length in exons of annotation was increased by 74% over our previously published annotation and includes 546 protein-coding genes and 234 pseudogenes. Thirty-two potential protein-coding annotations are partial copies of other genes, and may represent duplications on all evolutionary path to change or loss of function. We also identified 31 non-protein-coding transcripts, including 16 possible antisense RNAs. By extrapolation, we estimate the human genome contains 29,000-36,000 protein-coding genes, 21,300 pseudogenes, and 1500 antisense RNAs. We suggest that our revised annotation criteria provide a paradigm for future annotation of the human genome.