EasyGene--a prokaryotic gene finder that ranks ORFs by statistical significance.

EasyGene--a prokaryotic gene finder that ranks ORFs by statistical significance.
复制标题

DOI:
10.1186/1471-2105-4-21
复制
发表时间:
2003-06-03
期刊:
影响因子:
3
通讯作者:
Krogh A
Krogh A
中科院分区:
生物学4区
文献类型:
--
作者:
Larsen TS;Krogh A

文献摘要

参考文献

被引文献

相似文献

与序列分析的其他领域相反,一个假定基因的统计显著性度量还没有被设计出来,以帮助在原核生物基因组中从大量随机开放阅读框(orf)中区分真正的基因。因此,许多基因组有太多的短orf作为基因注释。在本文中,我们提出了一种新的自动基因查找方法,EasyGene,它估计预测基因的统计显著性。基因发现者是基于隐马尔可夫模型(HMM),自动估计一个新的基因组。利用Swiss-Prot中的相似性扩展,从基因组中自动提取高质量的基因训练集并用于估计HMM。然后用HMM对假定的基因进行评分,并根据评分和ORF的长度计算统计显著性。ORF的统计显著性度量是在相同或更好的显著性水平下随机序列的一个百万碱基中ORF的预期数量,其中随机序列在三阶马尔可夫链的意义上具有与基因组相同的统计量。结果是一个灵活的基因发现者,其整体性能匹配或超过其他方法。从基因组或组群的原始输入到具有重要意义的假定基因列表的整个计算机处理流程都是自动化的,这使得将EasyGene应用于新测序的生物体变得容易。EasyGene与预训练的模型可以访问。
Contrary to other areas of sequence analysis, a measure of statistical significance of a putative gene has not been devised to help in discriminating real genes from the masses of random Open Reading Frames (ORFs) in prokaryotic genomes. Therefore, many genomes have too many short ORFs annotated as genes. In this paper, we present a new automated gene-finding method, EasyGene, which estimates the statistical significance of a predicted gene. The gene finder is based on a hidden Markov model (HMM) that is automatically estimated for a new genome. Using extensions of similarities in Swiss-Prot, a high quality training set of genes is automatically extracted from the genome and used to estimate the HMM. Putative genes are then scored with the HMM, and based on score and length of an ORF, the statistical significance is calculated. The measure of statistical significance for an ORF is the expected number of ORFs in one megabase of random sequence at the same significance level or better, where the random sequence has the same statistics as the genome in the sense of a third order Markov chain. The result is a flexible gene finder whose overall performance matches or exceeds other methods. The entire pipeline of computer processing from the raw input of a genome or set of contigs to a list of putative genes with significance is automated, making it easy to apply EasyGene to newly sequenced organisms. EasyGene with pre-trained models can be accessed at .
DOI: 10.1093/nar/12.1part2.539
发表时间: 1984-01-01
影响因子: 14.9
作者:
GRIBSKOV, M;DEVEREUX, J;BURGESS, RR
通讯作者: BURGESS, RR
DOI: 10.1093/bioinformatics/17.12.1123
发表时间: 2001-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Suzek, BE;Ermolaeva, MD;Salzberg, SL
通讯作者: Salzberg, SL
DOI: 10.1093/dnares/6.2.83
发表时间: 1999-04-30
期刊: DNA research : an international journal for rapid publication of reports on genes and genomes
影响因子: --
作者:
Kawarabayasi, Y;Hino, Y;Kikuchi, H
通讯作者: Kikuchi, H
DOI: 10.1093/nar/29.12.2607
发表时间: 2001-06-15
影响因子: 14.9
作者:
Besemer, J;Lomsadze, A;Borodovsky, M
通讯作者: Borodovsky, M
DOI: 10.1093/nar/27.19.3911
发表时间: 1999-10-01
影响因子: 14.9
作者:
Besemer, J;Borodovsky, M
通讯作者: Borodovsky, M