Modeling and predicting transcriptional units of Escherichia coli genes using hidden Markov models

Modeling and predicting transcriptional units of Escherichia coli genes using hidden Markov models
复制标题

DOI:
10.1093/bioinformatics/15.12.987
复制
发表时间:
1999-12-01
期刊:
影响因子:
5.8
通讯作者:
Nakai, K
Nakai, K
中科院分区:
生物学3区
文献类型:
--
作者:
Yada, T;Nakao, M;Nakai, K

文献摘要

被引文献

相似文献

动机:隐马尔可夫模型(HMM)是一种有价值的基因发现技术,特别是因为它的灵活性,使包括各种序列特征。利用核糖体结合位点(RBS)的特性,提高了起始密码子识别的准确性。结果:首先,通过引入“典型”、“非典型”和“阴性”模型,提高了编码序列的预测精度(假阳性)类以及RBS及其下游间隔区的模型:对204例实验证实的CDS进行客观检验,准确预测的灵敏度达到90.2%。根据CDS的预测结果,预测了启动子和终止子的位置。我们的模型可以准确识别390个已知转录单位的60%。因此,这个预测问题的准确性和重要性远非微不足道。我们想把这个问题作为生物信息学的一个开放主题,因为正在进行或计划中的测序后项目将产生大量数据用于未来的改进。
Motivation: The hidden Markov model (HMM) is a valuable technique for gene-finding, especially because its flexibility enables the inclusion of various sequence features. Recent programs for bacterial gene-finding include the information of ribosomal binding site (RBS) to improve the recognition accuracy of the start codon, using this feature. We report here our attempt to extend the model into the total transcriptional unit, enabling the prediction of operon structures.Results: First, we improved the prediction accuracy of coding sequences (CDSs) by employing the models of 'typical', 'atypical' and 'negative (false-positive)' classes as well as the models of RBS and its downstream spacer: The sensitivity of exactly predicting the 204 experimentally confirmed CDSs reached 90.2% in an objective test. Based on the prediction result of CDSs, the positions of the promoters and terminators were predicted. Our model could exactly recognize 60% of 390 known transcriptional units. Thus, the accuracy and significance of this prediction problem is far from trivial. We would like to propose this problem as an open theme in bioinformatics because the ongoing or planned post-sequencing projects will produce much data for future improvements.