Gene prediction in novel fungal genomes using an ab initio algorithm with unsupervised training

Gene prediction in novel fungal genomes using an ab initio algorithm with unsupervised training
复制标题

DOI:
10.1101/gr.081612.108
复制
发表时间:
2008-12-01
期刊:
影响因子:
7
通讯作者:
Borodovsky, Mark
Borodovsky, Mark
中科院分区:
生物学1区
文献类型:
--
作者:
Ter-Hovhannisyan, Vardges;Lomsadze, Alexandre;Borodovsky, Mark

文献摘要

被引文献

相似文献

我们描述了一种新的从头算起算法,GeneMark-ES版本2,它识别真菌基因组中的蛋白质编码基因。该算法不需要预定的训练集来估计潜在的隐马尔可夫模型(HMM)的参数。取而代之的是,正在讨论的匿名基因组序列被用作迭代的无监督训练的输入。该算法扩展了我们之前开发的方法,在拟南芥、秀丽线虫和黑腹果蝇的基因组上进行了测试。为了更好地反映真菌基因组织的特征,我们增强了内含子模型,以适应带和不带分支点的序列。这种设计使算法能够同样好地适用于具有子囊菌门、担子菌门和接合菌门中所见的各种剪接机制的物种。在自我训练后,内含子模型分几个步骤启动,以达到其全部复杂性。我们证明,算法的准确性,无论是在外显子和整个基因水平上,都好于使用监督训练的基因搜索器的准确性。将新方法应用于已知的真菌基因组,表明比现有的注释有了实质性的改进。通过消除构建全面训练集的必要努力,新算法可以简化和加快大量真菌基因组测序项目的注释过程。
We describe a new ab initio algorithm, GeneMark-ES version 2, that identifies protein-coding genes in fungal genomes. The algorithm does not require a predetermined training set to estimate parameters of the underlying hidden Markov model (HMM). Instead, the anonymous genomic sequence in question is used as an input for iterative unsupervised training. The algorithm extends our previously developed method tested on genomes of Arabidopsis thaliana, Caenorhabditis elegans, and Drosophila melanogaster. To better reflect features of fungal gene organization, we enhanced the intron submodel to accommodate sequences with and without branch point sites. This design enables the algorithm to work equally well for species with the kinds of variations in splicing mechanisms seen in the fungal phyla Ascomycota, Basidiomycota, and Zygomycota. Upon self-training, the intron submodel switches on in several steps to reach its full complexity. We demonstrate that the algorithm accuracy, both at the exon and the whole gene level, is favorably compared to the accuracy of gene finders that employ supervised training. Application of the new method to known fungal genomes indicates substantial improvement over existing annotations. By eliminating the effort necessary to build comprehensive training sets, the new algorithm can streamline and accelerate the process of annotation in a large number of fungal genome sequencing projects.