MED: a new non-supervised gene prediction algorithm for bacterial and archaeal genomes.

MED: a new non-supervised gene prediction algorithm for bacterial and archaeal genomes.
复制标题

MED:一种新的细菌和古细菌基因组无监督基因预测算法。

DOI:
10.1186/1471-2105-8-97
复制
发表时间:
2007-03-16
期刊:
影响因子:
3
通讯作者:
She ZS
She ZS
中科院分区:
生物学4区
文献类型:
--
作者:
Zhu H;Hu GQ;Yang YF;Wang J;She ZS

文献摘要

参考文献

被引文献

相似文献

尽管在细菌和古生菌基因的计算预测方面取得了显著的成功,但缺乏对原核基因结构的全面了解阻碍了进一步阐明基因组之间的差异。开发新的从头算算法不仅可以准确地预测基因,而且还可以促进原核生物基因组的比较研究,这仍然是一个有趣的问题。本文提出了一种基于蛋白质编码开放阅读框(orf)和翻译起始位点(TISs)综合统计模型的原核基因查找算法。前者基于编码DNA序列的语言“熵密度曲线”(EDP)模型,后者包含与翻译起始相关的几个特征。它们被组合成一个所谓的多元熵距离(MED)算法,MED 2.0,它在迭代程序中包含了几种策略。迭代使我们能够开发一个无监督的学习过程,并在进行基因预测之前获得一组基因结构的基因组特异性参数。广泛的实验结果表明,与目前最好的原核基因查找器相比,MED 2.0在5‘和3’端匹配的基因预测方面都取得了具有竞争力的高性能。对于富含gc的基因组和古细菌基因组,MED 2.0的优势尤为明显。此外,MED 2.0给出的基因组特异性参数与目前对原核生物基因组的理解相吻合,可以作为比较基因组研究的工具。特别是,MED 2.0揭示了古细菌基因组中不同的翻译起始机制,与现有的基因发现器和当前的GenBank注释相比,可以更准确地预测TISs。
Despite a remarkable success in the computational prediction of genes in Bacteria and Archaea, a lack of comprehensive understanding of prokaryotic gene structures prevents from further elucidation of differences among genomes. It continues to be interesting to develop new ab initio algorithms which not only accurately predict genes, but also facilitate comparative studies of prokaryotic genomes. This paper describes a new prokaryotic genefinding algorithm based on a comprehensive statistical model of protein coding Open Reading Frames (ORFs) and Translation Initiation Sites (TISs). The former is based on a linguistic "Entropy Density Profile" (EDP) model of coding DNA sequence and the latter comprises several relevant features related to the translation initiation. They are combined to form a so-called Multivariate Entropy Distance (MED) algorithm, MED 2.0, that incorporates several strategies in the iterative program. The iterations enable us to develop a non-supervised learning process and to obtain a set of genome-specific parameters for the gene structure, before making the prediction of genes. Results of extensive tests show that MED 2.0 achieves a competitive high performance in the gene prediction for both 5' and 3' end matches, compared to the current best prokaryotic gene finders. The advantage of the MED 2.0 is particularly evident for GC-rich genomes and archaeal genomes. Furthermore, the genome-specific parameters given by MED 2.0 match with the current understanding of prokaryotic genomes and may serve as tools for comparative genomic studies. In particular, MED 2.0 is shown to reveal divergent translation initiation mechanisms in archaeal genomes while making a more accurate prediction of TISs compared to the existing gene finders and the current GenBank annotation.
DOI: 10.1186/1471-2105-4-21
发表时间: 2003-06-03
期刊: BMC bioinformatics
影响因子: 3
作者:
Larsen TS;Krogh A
通讯作者: Krogh A
DOI: 10.1142/s0219720004000624
发表时间: 2004-06-01
影响因子: 1
作者:
Ouyang, Zhengqing;Zhu, Huaiqiu;She, Zhen-Su
通讯作者: She, Zhen-Su
DOI: 10.1186/1471-2105-3-5
发表时间: 2002
期刊: BMC bioinformatics
影响因子: 3
作者:
Bocs S;Danchin A;Médigue C
通讯作者: Médigue C
DOI: 10.1038/ng0393-266
发表时间: 1993-03-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
GISH, W;STATES, DJ
通讯作者: STATES, DJ
DOI: 10.1093/bioinformatics/17.12.1123
发表时间: 2001-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Suzek, BE;Ermolaeva, MD;Salzberg, SL
通讯作者: Salzberg, SL