Accuracy improvement for identifying translation initiation sites in microbial genomes

Accuracy improvement for identifying translation initiation sites in microbial genomes
复制标题

DOI:
10.1093/bioinformatics/bth390
复制
发表时间:
2004-12-12
期刊:
影响因子:
5.8
通讯作者:
She, ZS
She, ZS
中科院分区:
生物学3区
文献类型:
--
作者:
Zhu, HQ;Hu, GQ;She, ZS

文献摘要

被引文献

相似文献

动机:目前微生物基因组中的计算基因识别方法对已验证的翻译终止位点(3‘端)有很高的预测精度,但对翻译起始位点(TIS,5’端)的预测精度要低得多。后者对于分析和理解基因的假定蛋白质和翻译的调控机制是很重要的。提高TIS预测的准确性是有待解决的问题之一。结果:本文建立了一个描述原核基因TIS的四组分统计模型。该模型包含了一些具有生物学意义的特征,包括基因翻译终止位点与TIS之间的相关性、起始密码子周围的序列含量、与核糖体结合位点相关的共识信号的序列含量以及TIS与上游共识信号之间的相关性。构建了一个完全无监督的训练系统,它以任何基因搜索者的一组带注释的编码开放阅读框架(ORF)作为输入,并给出一组生物体特有的参数(没有任何先验知识或经验常数和公式)作为输出。新算法在一组可靠的大肠杆菌和枯草芽孢杆菌基因数据集上进行了测试。MED-START可以正确预测195个实验确认的大肠杆菌基因中95.4%的起始点,以及58个可靠的枯草杆菌基因中96.6%的起始点。测试结果表明,该算法对较可靠的数据集具有较高的准确率,且对基因长度的变化具有较强的鲁棒性。MED-START可用作基因搜索器的后处理器。经过我们的程序处理后,基因发现系统对基因起始预测的改进是显著的,例如对于854个大肠杆菌验证基因,MED 1.0预测TIS的准确率从61.7%提高到91.5%,而Glimmer 2.02对相同数据集的预测准确率从63.2%提高到92.0%。这些结果表明,我们的算法是识别原核基因组TIS的最准确的方法之一。
Motivation: At present the computational gene identification methods in microbial genomes have a high prediction accuracy of verified translation termination site (3' end), but a much lower accuracy of the translation initiation site (TIS, 5' end). The latter is important to the analysis and the understanding of the putative protein of a gene and the regulatory machinery of the translation. Improving the accuracy of prediction of TIS is one of the remaining open problems.Results: In this paper, we develop a four-component statistical model to describe the TIS of prokaryotic genes. The model incorporates several features with biological meanings, including the correlation between translation termination site and TIS of genes, the sequence content around the start codon; the sequence content of the consensus signal related to ribosomal binding sites (RBSs), and the correlation between TIS and the upstream consensus signal. An entirely non-supervised training system is constructed, which takes as input a set of annotated coding open reading frames (ORFs) by any gene finder, and gives as output a set of organism-specific parameters (without any prior knowledge or empirical constants and formulas). The novel algorithm is tested on a set of reliable datasets of genes from Escherichia coli and Bacillus subtillis. MED-Start may correctly predict 95.4% of the start sites of 195 experimentally confirmed E.coli genes, 96.6% of 58 reliable B.subtillis genes. Moreover, the test results indicate that the algorithm gives higher accuracy for more reliable datasets, and is robust to the variation of gene length. MED-Start may be used as a postprocessor for a gene finder. After processing by our program, the improvement of gene start prediction of gene finder system is remarkable, e.g. the accuracy of TIS predicted by MED 1.0 increases from 61.7 to 91.5% for 854 E.coli verified genes, while that by GLIMMER 2.02 increases from 63.2 to 92.0% for the same dataset. These results show that our algorithm is one of the most accurate methods to identify TIS of prokaryotic genomes.