Missing genes in the annotation of prokaryotic genomes.

Missing genes in the annotation of prokaryotic genomes.
复制标题

DOI:
10.1186/1471-2105-11-131
复制
发表时间:
2010-03-15
期刊:
影响因子:
3
通讯作者:
Setubal JC
Setubal JC
中科院分区:
生物学4区
文献类型:
--
作者:
Warren AS;Archuleta J;Feng WC;Setubal JC

文献摘要

参考文献

被引文献

相似文献

原核基因组中的蛋白质编码基因检测被认为比含有内含子的真核基因组中的蛋白质编码基因检测简单得多。然而,有报道称,原核基因查找程序在处理小基因时存在问题(要么过度预测,要么预测不足)。因此,问题是当前的基因组注释是否系统性缺失了小基因。我们开发了一种高性能计算方法来研究这个问题。在这种方法中,我们比较所有全测序原核复制子中大于或等于 33 个氨基酸的所有 ORF。基于这一比较,并使用要求不同基因组中保守 ORF 之间具有最小分类多样性的保守标准,我们发现了当前基因组注释中缺失的 1,153 个候选基因。这些缺失基因仅彼此相似,与公共数据库中的基因序列没有任何很强的相似性,这意味着这些 ORF 属于缺失基因家族。我们还发现了 38,895 个基因间 ORF,通过与当前注释基因(我们称之为缺失注释)的相似性,很容易将其识别为推定基因。发现的绝大多数缺失基因都很小(小于 100 个氨基酸)。将选定的示例与 GeneMark、EasyGene 和 Glimmer 预测进行比较,得出的证据表明其中一些基因正在逃避这些程序的检测。原核基因查找器和原核基因组注释需要改进,以准确预测小基因。由于用于确定 ORF 是否对应于真实基因的保守标准,发现的缺失基因家族的数量可能是实际数量的下限。
Protein-coding gene detection in prokaryotic genomes is considered a much simpler problem than in intron-containing eukaryotic genomes. However there have been reports that prokaryotic gene finder programs have problems with small genes (either over-predicting or under-predicting). Therefore the question arises as to whether current genome annotations have systematically missing, small genes. We have developed a high-performance computing methodology to investigate this problem. In this methodology we compare all ORFs larger than or equal to 33 aa from all fully-sequenced prokaryotic replicons. Based on that comparison, and using conservative criteria requiring a minimum taxonomic diversity between conserved ORFs in different genomes, we have discovered 1,153 candidate genes that are missing from current genome annotations. These missing genes are similar only to each other and do not have any strong similarity to gene sequences in public databases, with the implication that these ORFs belong to missing gene families. We also uncovered 38,895 intergenic ORFs, readily identified as putative genes by similarity to currently annotated genes (we call these absent annotations). The vast majority of the missing genes found are small (less than 100 aa). A comparison of select examples with GeneMark, EasyGene and Glimmer predictions yields evidence that some of these genes are escaping detection by these programs. Prokaryotic gene finders and prokaryotic genome annotations require improvement for accurate prediction of small genes. The number of missing gene families found is likely a lower bound on the actual number, due to the conservative criteria used to determine whether an ORF corresponds to a real gene.
DOI: 10.1186/1471-2105-4-21
发表时间: 2003-06-03
期刊: BMC bioinformatics
影响因子: 3
作者:
Larsen TS;Krogh A
通讯作者: Krogh A
DOI: 10.1093/nar/gkn664
发表时间: 2009-01
影响因子: 14.9
作者:
UniProt Consortium
通讯作者: UniProt Consortium
DOI: 10.1073/pnas.0409727102
发表时间: 2005-02-15
影响因子: 11.1
作者:
Konstantinidis, KT;Tiedje, JM
通讯作者: Tiedje, JM
DOI: 10.1128/jb.187.18.6258-6264.2005
发表时间: 2005-09-01
影响因子: 3.2
作者:
Konstantinidis, KT;Tiedje, JM
通讯作者: Tiedje, JM
DOI: 10.1128/jb.00872-09
发表时间: 2010-01-01
影响因子: 3.2
作者:
Hemm, Matthew R.;Paul, Brian J.;Storz, Gisela
通讯作者: Storz, Gisela