A revised annotation and comparative analysis of Helicobacter pylori genomes

A revised annotation and comparative analysis of Helicobacter pylori genomes
复制标题

DOI:
10.1093/nar/gkg250
复制
发表时间:
2003-03-15
影响因子:
14.9
通讯作者:
Moszer, I
Moszer, I
中科院分区:
生物学2区
文献类型:
--
作者:
Boneca, IG;de Reuse, H;Moszer, I

文献摘要

被引文献

相似文献

目前正在产生大量的基因组信息。因此,生物学家需要结构化的、详尽的和比较的数据库。幽门螺杆菌基因数据库(http://genolist.pasteur.fr/PyloriGene))是为了响应这些需求而开发的,通过整合和连接在两种不同的幽门螺杆菌菌株测序过程中产生的信息。这导致了对一般性注释共识的需要,因为在某些情况下,这两个菌株的物理和功能注释显著不同。建立了一个修订的功能分类系统,以适应现有数据,并使其能够将编码序列(CDS)分类为几个功能类别,以协调CDS分类。两个完整基因组的注释根据新的数据进行了修改,使我们能够将假设蛋白质的百分比从类似的40%降低到33%。这导致重新分配了108个CDS的职能(相当于所有CDS的7%)。有趣的是,只有13%的CDS(1658个CDS中的222个)的功能被注释,这是直接对H.Pylori基因进行研究的结果。最后,对两个已发表的基因组进行比较,发现相应的(同源的)CDS之间存在显著的大小差异。这些大小的变异大多是由自然的多态引起的,尽管也发现了其他的变异来源,如假基因、可能受滑链错配机制调控的新基因或移码。其中113个差异是由于不同的起始密码子分配,这是构建物理注释时的一个常见问题。
Huge amounts of genomic information are currently being generated. Therefore, biologists require structured, exhaustive and comparative databases. The PyloriGene database (http://genolist.pasteur.fr/PyloriGene) was developed to respond to these needs, by integrating and connecting the information generated during the sequencing of two distinct strains of Helicobacter pylori. This led to the need for a general annotation consensus, as the physical and functional annotations of the two strains differed significantly in some cases. A revised functional classification system was created to accommodate the existing data and to make it possible to classify coding sequences (CDS) into several functional categories to harmonize CDS classification. The annotation of the two complete genomes was revised in the light of new data, allowing us to reduce the percentage of hypothetical proteins from similar to40 to 33%. This resulted in the reassignment of functions for 108 CDS (similar to7% of all CDS). Interestingly, the functions of only similar to13% of CDS (222 out of 1658 CDS) were annotated as a result of work done directly on H.pylori genes. Finally, comparison of the two published genomes revealed a significant amount of size variation between corresponding (orthologous) CDS. Most of these size variations were due to natural polymorphisms, although other sources of variation were identified, such as pseudogenes, new genes potentially regulated by slipped-strand mispairing mechanism, or frame-shifts. 113 of these differences were due to different start codon assignments, a common problem when constructing physical annotations.