Operon Prediction Based On an Iterative Self-learning Algorithm

Operon Prediction Based On an Iterative Self-learning Algorithm
复制标题

基于迭代自学习算法的操纵子预测

DOI:
--
复制
发表时间:
2011
影响因子:
0.3
通讯作者:
Zhu Huai-Qiu
Zhu Huai-Qiu
中科院分区:
生物学4区
文献类型:
--
作者:
Wu Wen-Qi;Zheng Xiao-Bin;Liu Yong-Chu;Tang Kai;Zhu Huai-Qiu

文献摘要

相似文献

操纵子作为原核基因组中基因的一种特定功能组织,包含一组在相应调控信号控制下的相邻基因,并以转录单位的形式表达,研究发现一个操纵子中的基因通常倾向于具有相关功能,因此,研究操纵子的结构对了解基因的功能和调控网络具有重要意义,然而,由于目前操纵子的数据获取受到诸如原核转录组学的实验的限制,注释新测序的基因组中的操纵子的计算方法迄今为止已经是操纵子数据的主要来源,并且将继续是重要的使命。在过去的十年中,已经提出了一组操纵子预测的计算方法,针对操纵子预测的瓶颈问题,提出了一种不依赖于已知操纵子数据集的迭代自学习算法,该算法基于概率模型,利用基因距离等特征,基因表达的调控信号和功能注释,如COG。与实验操纵子数据进行比较的测试结果表明,该算法在没有任何训练集的情况下可以达到最佳的准确率。此外,这种自学习算法上级在已知操纵子的任何物种上训练的算法。相应地,该算法可以应用于任何新测序的基因组。2此外,细菌和古细菌的比较分析增强了对操纵子的普遍和基因组特异性特征的认识。
As a specific functional organization of genes in prokaryotic genomes,operon contains a set of adjacent genes under the control of the corresponding regulatory signals,and is expressed as the transcript unit.It has been found that genes in an operon usually tend to have related functions,or belong to the same pathway in cell.Therefore the study of operon structure is significant to understand the gene functions and regulatory networks for prokaryotes.However with the current limitation of data acquisition of operons verified by experiments such as prokaryotic transcriptomics,computation methods to annotate the operons in a newly sequenced genome have so far been the major source of operon data,and will continue to be an important mission.Over the past decade,a set of computational approaches to operon prediction have been proposed,however mainly based on experimental operons as their training sets.Nevertheless the lack of experimental operon dataset has been the bottleneck of operon prediction.The authors employ an iterative self-learning algorithm which is independent of training set with known operon dataset.The algorithm develops based on a probabilistic model using features including gene distance,regulation signals of gene expression and functional annotation such as COG.The test result compared against the experimental operon data indicates that the algorithm can reach the best accuracy without any training set.Besides,this self-learning algorithm is superior to the algorithm trained on any species with known operons.Accordingly,the algorithm can be applied to any newly sequenced genome.Moreover,comparative analysis of bacteria and archaea enhances the knowledge of universal and genome specific features of operons.