Genome annotation assessment in Drosophila melanogaster

Genome annotation assessment in Drosophila melanogaster
复制标题

DOI:
10.1101/gr.10.4.483
复制
发表时间:
2000-04-01
期刊:
影响因子:
7
通讯作者:
Lewis, SE
Lewis, SE
中科院分区:
生物学1区
文献类型:
--
作者:
Reese, MG;Hartzell, G;Lewis, SE

文献摘要

被引文献

相似文献

用于自动化基因组注释的计算方法对于我们社区充分利用正在生成和发布的大量基因组序列的能力至关重要。为了探索这些自动化特征预测工具在高等生物基因组中的准确性,我们评估了它们在果蝇Adh区域的大型,特征良好的序列重叠群上的性能。这项实验被称为基因组注释评估项目(GASP),于1999年5月启动。12个小组,应用最先进的工具,贡献了预测功能,包括基因结构,蛋白质同源性,启动子sires和重复元件。我们使用两种标准评估这些预测,一种是基于以前未发布的高质量全长cDNA序列,另一种是基于一组果蝇专家对该区域进行深入研究所产生的一组注释。虽然这些标准集仅近似于该区域中未知的特征分布,但我们相信,当在上下文中考虑时,基于它们的评估结果是有意义的。研究结果在1999年8月的分子生物学智能系统会议(ISMB-99)上作为教程发表。大多数基因发现者正确鉴定了该区域超过95%的编码核苷酸,并预测了>40%的基因的正确内含子/外显子结构。基于同源性的注释技术识别并关联了该区域中几乎一半的基因的功能;其余的仅通过从头算技术识别。该实验还提出了启动子预测技术在一个大的连续区域中的大量基因的第一次评估。我们发现,启动子预测因子的高假阳性率使他们的预测难以使用。整合基因发现和cDNA/EST比对与启动子预测减少了假阳性分类的数量,但发现不到三分之一的启动子在该地区。我们认为,通过建立评估基因组注释的标准,并通过评估现有的自动化基因组注释工具的性能,该实验建立了一个基线,有助于正在进行的大规模注释项目的价值,并应指导基因组信息学的进一步研究。
Computational methods for automated genome annotation are critical to our community's ability to make full use of the large volume of genomic sequence being generated and released. To explore the accuracy of these automated feature prediction tools in the genomes of higher organisms, we evaluated their performance on a large, well-characterized sequence contig from the Adh region of Drosophila melanogaster. This experiment, known as the Genome Annotation Assessment Project (GASP), was launched in May 1999. Twelve groups, applying state-of-the-art tools, contributed predictions for features including gene structure, protein homologies, promoter sires, and repeat elements. We evaluated these predictions using two standards, one based on previously unreleased high-quality full-length cDNA sequences and a second based on the set of annotations generated as part of an in-depth study of the region by a group of Drosophila experts. Although these standard sets only approximate the unknown distribution of Features in this region, we believe that when taken in context the results of an evaluation based on them are meaningful. The results were presented as a tutorial at the conference on intelligent Systems in Molecular Biology (ISMB-99) in August 1999. Over 95% of the coding nucleotides in the region were correctly identified by the majority of the gene finders, and the correct intron/exon structures were predicted For >40% of the genes. Homology-based annotation techniques recognized and associated functions with almost half of the genes in the region; the remainder were only identified by the ab initio techniques. This experiment also presents the first assessment of promoter prediction techniques for a significant number of genes in a large contiguous region. We discovered that the promoter predictors' high false-positive rates make their predictions difficult to use. Integrating gene Finding and cDNA/EST alignments with promoter predictions decreases the number of false-positive classifications but discovers less than one-third of the promoters in the region. We believe that by establishing standards for evaluating genomic annotations and by assessing the performance of existing automated genome annotation tools, this experiment establishes a baseline that contributes to the value of ongoing large-scale annotation projects and should guide further research in genome informatics.