Full-length messenger RNA sequences greatly improve genome annotation.

Full-length messenger RNA sequences greatly improve genome annotation.
复制标题

DOI:
10.1186/gb-2002-3-6-research0029
复制
发表时间:
2002
期刊:
影响因子:
12.3
通讯作者:
Salzberg SL
Salzberg SL
中科院分区:
生物学1区
文献类型:
--
作者:
Haas BJ;Volfovsky N;Town CD;Troukhan M;Alexandrov N;Feldmann KA;Flavell RB;White O;Salzberg SL

文献摘要

被引文献

相似文献

真核生物基因组注释是一项复杂的工作,需要整合来自多个常常相互矛盾的数据源的证据。随着现在可用的基因组序列数据量不断增加,准确识别大量基因的方法变得迫切需要。为了创建一组高质量的基因模型,我们利用来自拟南芥的5000个全长基因转录本的序列对其基因组进行重新注释。我们已将这些转录本定位到它们确切的染色体位置,并利用比对程序创建了基因模型,为该生物提供了一个参考集。 大约35%的转录本表明先前注释的基因需要修改,5%的转录本代表新发现的基因。我们还发现多个转录起始位点似乎比先前所知的更为常见,并且我们报道了许多选择性mRNA剪接的案例。我们对不同的比对软件进行了比较,并分析了转录本数据如何改进先前发表的注释。 我们的结果表明,对大量全长转录本进行测序,然后进行计算机定位,极大地提高了对真核生物基因完整外显子结构的识别。此外,我们能够在基因的非翻译区发现许多内含子。
Annotation of eukaryotic genomes is a complex endeavor that requires the integration of evidence from multiple, often contradictory, sources. With the ever-increasing amount of genome sequence data now available, methods for accurate identification of large numbers of genes have become urgently needed. In an effort to create a set of very high-quality gene models, we used the sequence of 5,000 full-length gene transcripts from Arabidopsis to re-annotate its genome. We have mapped these transcripts to their exact chromosomal locations and, using alignment programs, have created gene models that provide a reference set for this organism. Approximately 35% of the transcripts indicated that previously annotated genes needed modification, and 5% of the transcripts represented newly discovered genes. We also discovered that multiple transcription initiation sites appear to be much more common than previously known, and we report numerous cases of alternative mRNA splicing. We include a comparison of different alignment software and an analysis of how the transcript data improved the previously published annotation. Our results demonstrate that sequencing of large numbers of full-length transcripts followed by computational mapping greatly improves identification of the complete exon structures of eukaryotic genes. In addition, we are able to find numerous introns in the untranslated regions of the genes.