Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies

Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies
复制标题

DOI:
10.1093/nar/gkg770
复制
发表时间:
2003-10-01
影响因子:
14.9
通讯作者:
White, O
White, O
中科院分区:
生物学2区
文献类型:
--
作者:
Haas, BJ;Delcher, AL;White, O

文献摘要

被引文献

相似文献

表达序列数据与基因组序列的拼接比对已被证明是对真核基因组中基因进行全面注释的关键工具。开发了一种新的算法来将重叠的转录本比对簇(EST和全长cDNA)组装成最大比对组合,从而全面整合所有可用的转录本数据并捕获细微的剪接变异。用这种方法确定的完整和部分基因结构被用于改进拟南芥基因组注释研究所(TIGR Release v.4.0)。这些比对组件允许对几个新基因和>1000选择性剪接变异进行自动化建模,以及对27,000个带注释的蛋白质编码基因中近一半的基因进行更新(包括UTR注释)。描述了拼接比对程序(PASA)工具的算法,以及对拟南芥基因注释进行自动更新的结果。
The spliced alignment of expressed sequence data to genomic sequence has proven a key tool in the comprehensive annotation of genes in eukaryotic genomes. A novel algorithm was developed to assemble clusters of overlapping transcript alignments (ESTs and full-length cDNAs) into maximal alignment assemblies, thereby comprehensively incorporating all available transcript data and capturing subtle splicing variations. Complete and partial gene structures identified by this method were used to improve The Institute for Genomic Research Arabidopsis genome annotation (TIGR release v.4.0). The alignment assemblies permitted the automated modeling of several novel genes and >1000 alternative splicing variations as well as updates (including UTR annotations) to nearly half of the similar to27 000 annotated protein coding genes. The algorithm of the Program to Assemble Spliced Alignments (PASA) tool is described, as well as the results of automated updates to Arabidopsis gene annotations.