Re-annotation of genome microbial coding-sequences: finding new genes and inaccurately annotated genes.

Re-annotation of genome microbial coding-sequences: finding new genes and inaccurately annotated genes.
复制标题

DOI:
10.1186/1471-2105-3-5
复制
发表时间:
2002
期刊:
影响因子:
3
通讯作者:
Médigue C
Médigue C
中科院分区:
生物学4区
文献类型:
--
作者:
Bocs S;Danchin A;Médigue C

文献摘要

被引文献

相似文献

对任何新测序的细菌基因组的分析都从蛋白质编码基因的识别开始。尽管积累了多个完整的基因组序列,可以在注释过程中与其他生物体的近亲进行有用的比较,但准确的基因预测仍然相当困难。造成这种情况的一个主要原因是原核生物中基因紧密堆积,导致频繁重叠。因此,除非该方法中嵌入了适当的生物学知识(关于基因的结构),否则检测翻译起始位点和/或选择正确的编码区仍然很困难。我们开发了一种新程序,可以自动识别细菌基因组中具有生物学意义的候选基因。使用该工具分析了二十六个完整的原核生物基因组,并通过与现有注释进行比较来评估基因发现的准确性。这项分析表明,尽管基因组程序注释者付出了巨大的努力,但在测序项目框架内注释的一小部分但不可忽略的基因可能是部分不准确或明显错误的。此外,对几个假定的新基因的分析表明,正如预期的那样,许多短基因已经逃脱了注释。在大多数情况下,这些新基因揭示的移码可能是伪影或真正的移码。一些完全意想不到的新基因也已被发现。这使我们能够更完整地了解原核生物基因组。该程序的结果将逐步集成到 SWISS-PROT 参考数据库中。本研究中描述的结果表明,我们的程序在基因查找准确性方面非常令人满意。除少数情况外,我们的结果与个别作者提供的注释之间的差异可以由每个注释过程的性质或某些基因组的特定特征来解释。这强调,显然需要科学家之间的密切合作、定期更新和管理数据库中的发现,以减少基因组注释中的错误水平(从而减少通过集中数据库不幸传播的错误)。
Analysis of any newly sequenced bacterial genome starts with the identification of protein-coding genes. Despite the accumulation of multiple complete genome sequences, which provide useful comparisons with close relatives among other organisms during the annotation process, accurate gene prediction remains quite difficult. A major reason for this situation is that genes are tightly packed in prokaryotes, resulting in frequent overlap. Thus, detection of translation initiation sites and/or selection of the correct coding regions remain difficult unless appropriate biological knowledge (about the structure of a gene) is imbedded in the approach. We have developed a new program that automatically identifies biologically significant candidate genes in a bacterial genome. Twenty-six complete prokaryotic genomes were analyzed using this tool, and the accuracy of gene finding was assessed by comparison with existing annotations. This analysis revealed that, despite the enormous effort of genome program annotators, a small but not negligible number of genes annotated within the framework of sequencing projects are likely to be partially inaccurate or plainly wrong. Moreover, the analysis of several putative new genes shows that, as expected, many short genes have escaped annotation. In most cases, these new genes revealed frameshifts that could be either artifacts or genuine frameshifts. Some entirely unexpected new genes have also been identified. This allowed us to get a more complete picture of prokaryotic genomes. The results of this procedure are progressively integrated into the SWISS-PROT reference databank. The results described in the present study show that our procedure is very satisfactory in terms of gene finding accuracy. Except in few cases, discrepancies between our results and annotations provided by individual authors can be accounted for by the nature of each annotation process or by specific characteristics of some genomes. This stresses that close cooperation between scientists, regular update and curation of the findings in databases are clearly required to reduce the level of errors in genome annotation (and hence in reducing the unfortunate spreading of errors through centralized data libraries).