A 'PolyORFomic' analysis of prokaryote genomes using disabled-homology filtering reveals conserved but undiscovered short ORFs

A 'PolyORFomic' analysis of prokaryote genomes using disabled-homology filtering reveals conserved but undiscovered short ORFs
复制标题

DOI:
10.1016/j.jmb.2003.09.016
复制
发表时间:
2003-11-07
影响因子:
5.6
通讯作者:
Gerstein, M
Gerstein, M
中科院分区:
生物学2区
文献类型:
--
作者:
Harrison, PM;Carriero, N;Gerstein, M

文献摘要

被引文献

相似文献

原核生物基因注释由于遗传密码设计中自然产生的大量短开放阅读框(ORF)而变得复杂。历史上,许多假设的ORF被注释为微生物中的基因,通常具有任意长度阈值(例如大于100个密码子)。考虑到这些阈值的使用,在目前的原核生物基因组样本中,真正未被发现的短基因的范围是多少?为了严格评估具有同源性的短ORF的潜在注释不足,我们详尽地比较了polyORFome-64种原核生物(53种细菌和11种古生菌)加上出芽酵母中所有可能的ORF-本身和所有已知的蛋白质。我们的分析的新奇在于,首先,被认为是/之间的注释和未注释的ORF的序列比较,其次,一个两步禁用同源性过滤器被施加到搁置假定的假基因和假ORF。我们发现,未注释的同源短ORF(uhORF)对应于注释的原核生物蛋白质组的一个小的,但不可忽略的部分(0.5- 3.8%,取决于选择标准)。此外,禁用同源性过滤器表明,大约三分之一的uhORF对应于推定的假基因或假ORF。我们的分析表明,注释长度阈值的使用是不必要的,因为在微生物基因组中有可管理数量的短ORF同源性保守(没有禁用)。uhORF的数据可从http://pseudogene.org/polyo(C)2003 Elsevier Ltd.获得。保留所有权利。
Prokaryote gene annotation is complicated by large numbers of short open reading frames (ORFs) that arise naturally from genetic code design. Historically, many hypothetical ORFs have been annotated as genes in microbes, usually with an arbitrary length threshold (e.g. greater than 100 codons). Given the use of such thresholds, what is the extent of genuine undiscovered short genes in the current sampling of prokaryote genomes? To assess rigorously the potential under-annotation of short ORFs with homology, we exhaustively compared the polyORFome-all possible ORFs in 64 prokaryotes (53 bacteria and 11 archaea) plus budding yeast-to itself and to all known proteins. The novelty of our analysis is that, firstly, sequence comparisons to/between both annotated and un-annotated ORFs are considered, and secondly a two-step disabled-homology filter is applied to set aside putative pseudogenes and spurious ORFs. We find that un-annotated homologous short ORFs (uhORFs) correspond to a small but non-negligible fraction of the annotated prokaryote proteomes (0.5-3.8%, depending on selection criteria). Moreover, the disabled-homology filter indicates that about a third of uhORFs correspond to putative pseudogenes or spurious ORFs. Our analysis shows that the use of annotation length thresholds is unnecessary, as there are manageable numbers of short ORF homologies conserved (without disablements) across microbial genomes. Data on uhORFs are available from http://pseudogene.org/polyo (C) 2003 Elsevier Ltd. All rights reserved.