Gene identification in novel eukaryotic genomes by self-training algorithm.

Gene identification in novel eukaryotic genomes by self-training algorithm.
复制标题

通过自我训练算法在新型真核基因组中的基因鉴定。

DOI:
10.1093/nar/gki937
复制
发表时间:
2005
影响因子:
14.9
通讯作者:
Borodovsky M
Borodovsky M
中科院分区:
生物学2区
文献类型:
--
作者:
Lomsadze A;Ter-Hovhannisyan V;Chernoff YO;Borodovsky M

文献摘要

参考文献

被引文献

相似文献

寻找新的蛋白质编码基因是真核基因组测序项目最重要的目标之一。然而,新的真核基因组的基因组组织是多样的,针对先前研究的物种而改进的从头算基因寻找工具很少适合于在新基因组的DNA序列中有效地寻找基因。基于基因组DNA的cDNA和表达序列标签(EST)定位的基因识别方法,或基于与之密切相关的基因组序列比对的基因识别方法,要么依赖于丰富的cDNA和EST数据的存在,要么依赖于参考基因组的可用性。传统的统计从头算方法需要大量有效基因的训练集来估计基因模型参数。在实践中,直到新的基因组测序的相当晚的阶段,这些类型的数据中的任何一种都不可能有足够的量。然而,我们已经证明,在真核基因组中发现基因可以与直接从匿名基因组DNA进行统计模型估计并行进行。所提出的基因预测与模型参数估计并行的方法遵循迭代维特比训练的路径。基因组序列被标记为编码区和非编码区的几轮之后是模型参数估计的几轮。对模型参数的可能范围增加了几个动态变化的限制,以滤除算法初始步骤中的波动,这些波动可能会将迭代过程重新定向到参数空间中的生物相关点。对研究充分的真核基因组的测试表明,新方法的性能与传统方法相当或更好,在传统方法中,监督模型训练先于基因预测步骤。已经分析了几个新的基因组,并讨论了生物学上有趣的发现。因此,一种被认为只适用于原核基因组的自训练算法现在已经被开发用于从头开始真核基因识别。
Finding new protein-coding genes is one of the most important goals of eukaryotic genome sequencing projects. However, genomic organization of novel eukaryotic genomes is diverse and ab initio gene finding tools tuned up for previously studied species are rarely suitable for efficacious gene hunting in DNA sequences of a new genome. Gene identification methods based on cDNA and expressed sequence tag (EST) mapping to genomic DNA or those using alignments to closely related genomes rely either on existence of abundant cDNA and EST data and/or availability on reference genomes. Conventional statistical ab initio methods require large training sets of validated genes for estimating gene model parameters. In practice, neither one of these types of data may be available in sufficient amount until rather late stages of the novel genome sequencing. Nevertheless, we have shown that gene finding in eukaryotic genomes could be carried out in parallel with statistical models estimation directly from yet anonymous genomic DNA. The suggested method of parallelization of gene prediction with the model parameters estimation follows the path of the iterative Viterbi training. Rounds of genomic sequence labeling into coding and non-coding regions are followed by the rounds of model parameters estimation. Several dynamically changing restrictions on the possible range of model parameters are added to filter out fluctuations in the initial steps of the algorithm that could redirect the iteration process away from the biologically relevant point in parameter space. Tests on well-studied eukaryotic genomes have shown that the new method performs comparably or better than conventional methods where the supervised model training precedes the gene prediction step. Several novel genomes have been analyzed and biologically interesting findings are discussed. Thus, a self-training algorithm that had been assumed feasible only for prokaryotic genomes has now been developed for ab initio eukaryotic gene identification.
DOI: 10.1186/1471-2105-4-21
发表时间: 2003-06-03
期刊: BMC bioinformatics
影响因子: 3
作者:
Larsen TS;Krogh A
通讯作者: Krogh A
DOI: 10.1038/ng0393-266
发表时间: 1993-03-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
GISH, W;STATES, DJ
通讯作者: STATES, DJ
DOI: 10.1089/cmb.1998.5.307
发表时间: 1998-06-01
影响因子: 1.7
作者:
Laub, MT;Smith, DW
通讯作者: Smith, DW
DOI: 10.1093/bioinformatics/18.6.777
发表时间: 2002-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Morgenstern, B;Rinner, O;Mewes, HW
通讯作者: Mewes, HW
DOI: 10.1093/nar/29.12.2607
发表时间: 2001-06-15
影响因子: 14.9
作者:
Besemer, J;Lomsadze, A;Borodovsky, M
通讯作者: Borodovsky, M