A probabilistic, method for identifying start codons in bacterial genomes

A probabilistic, method for identifying start codons in bacterial genomes
复制标题

DOI:
10.1093/bioinformatics/17.12.1123
复制
发表时间:
2001-12-01
期刊:
影响因子:
5.8
通讯作者:
Salzberg, SL
Salzberg, SL
中科院分区:
生物学3区
文献类型:
--
作者:
Suzek, BE;Ermolaeva, MD;Salzberg, SL

文献摘要

被引文献

相似文献

随着基因组测序步伐的加快,对高度准确的基因预测系统的需求也在增长。用于鉴定原核基因组中的基因的计算系统具有98-99%或更高的灵敏度(Delcher等人,核酸研究,27,4636-4641,1999)。这些准确度数字是通过比较验证的终止密码子的位置与预测来计算的。然而,确定起始密码子预测的准确性更成问题,这是由于已经通过独立的非计算方法确认的起始位点的数量相对较少。尽管如此,基因发现者在预测基因的5'和3'端的确切基因边界方面的准确性对于微生物基因组注释是至关重要的,特别是考虑到有时在蛋白质编码区的5'端发现的重要信号传导信息。在本文中,我们提出了一种概率方法,以提高基因识别系统在寻找精确的翻译起始位点的准确性。新系统RBSfinder在一组经过验证的大肠杆菌基因上进行了测试,它提高了计算基因发现系统预测的起始位点位置的准确性,从67-77%到90%的正确率。
As the pace of genome sequencing has accelerated, the need for highly accurate gene prediction systems has grown. Computational systems for identifying genes in prokaryotic genomes have sensitivities of 98-99% or higher (Delcher et al., Nucleic Acids Res., 27, 4636-4641, 1999). These accuracy figures are calculated by comparing the locations of verified stop codons to the predictions. Determining the accuracy of start codon prediction is more problematic, however, due to the relatively small number of start sites that have been confirmed by independent, non-computational methods. Nonetheless, the accuracy of gene finders at predicting the exact gene boundaries at both the 5' and 3' ends of genes is of critical importance for microbial genome annotation, especially in light of the important signaling information that is sometimes found on the 5' end of a protein coding region. In this paper we propose a probabilistic method to improve the accuracy of gene identification systems at finding precise translation start sites. The new system, RBSfinder, is tested on a validated set of genes from Escherichia coli, for which it improves the accuracy of start site locations predicted by computational gene finding systems from the range 67-77% to 90% correct.