Recognition of prokaryotic promoters based on a novel variable-window Z-curve method.

Recognition of prokaryotic promoters based on a novel variable-window Z-curve method.
复制标题

基于新型可变窗口Z曲线方法的原核启动子识别

DOI:
10.1093/nar/gkr795
复制
发表时间:
2012-02
影响因子:
14.9
通讯作者:
Song K
Song K
中科院分区:
生物学2区
文献类型:
--
作者:
Song K

文献摘要

参考文献

被引文献

相似文献

转录是基因表达的第一步,也是大多数表达调控发生的步骤。虽然测序的原核基因组提供了丰富的信息,转录调控网络仍然是知之甚少,使用现有的基因组信息,主要是因为准确预测的启动子是困难的。为了提高启动子识别性能,提出了一种新的可变窗口Z曲线方法来提取原核启动子的一般特征。这些特征被用于通过偏最小二乘技术进行进一步分类。为了验证预测性能,所提出的方法被应用于预测两个代表性的原核模式生物(大肠杆菌和枯草芽孢杆菌)的启动子片段。根据所提出的方法的特征提取和选择能力,启动子预测的准确性显着提高了大多数现有的方法:对于E。对大肠杆菌σ70启动子、σ70启动子、已知σ因子启动子、非编码阴性样品的准确率分别为96.05%、90.44%、92.13%、92.50%;在枯草杆菌中,准确率为95.83%(已知σ因子启动子,编码阴性样品)和99.09%(已知σ因子启动子,非编码阴性样品)。此外,作为一种线性技术,所提出的方法的计算简单性使得它很容易在普通个人计算机甚至笔记本电脑上运行几分钟。更重要的是,不需要优化参数,因此,它是非常实用的预测其他物种的启动子没有任何先验知识或先验信息的统计性质的样品。
Transcription is the first step in gene expression, and it is the step at which most of the regulation of expression occurs. Although sequenced prokaryotic genomes provide a wealth of information, transcriptional regulatory networks are still poorly understood using the available genomic information, largely because accurate prediction of promoters is difficult. To improve promoter recognition performance, a novel variable-window Z-curve method is developed to extract general features of prokaryotic promoters. The features are used for further classification by the partial least squares technique. To verify the prediction performance, the proposed method is applied to predict promoter fragments of two representative prokaryotic model organisms (Escherichia coli and Bacillus subtilis). Depending on the feature extraction and selection power of the proposed method, the promoter prediction accuracies are improved markedly over most existing approaches: for E. coli, the accuracies are 96.05% (σ70 promoters, coding negative samples), 90.44% (σ70 promoters, non-coding negative samples), 92.13% (known sigma-factor promoters, coding negative samples), 92.50% (known sigma-factor promoters, non-coding negative samples), respectively; for B. subtilis, the accuracies are 95.83% (known sigma-factor promoters, coding negative samples) and 99.09% (known sigma-factor promoters, non-coding negative samples). Additionally, being a linear technique, the computational simplicity of the proposed method makes it easy to run in a matter of minutes on ordinary personal computers or even laptops. More importantly, there is no need to optimize parameters, so it is very practical for predicting other species promoters without any prior knowledge or prior information of the statistical properties of the samples.
DOI: 10.1266/ggs.84.425
发表时间: 2009-12-01
影响因子: 1.1
作者:
Askary, Amjad;Masoudi-Nejad, Ali;Purmasjedi, Malihe
通讯作者: Purmasjedi, Malihe
DOI: 10.1186/gb-2003-4-1-203
发表时间: 2003
期刊: Genome biology
影响因子: 12.3
作者:
Paget MS;Helmann JD
通讯作者: Helmann JD
DOI: 10.3233/isb-2009-0388
发表时间: 2009-01-01
期刊: In Silico Biology
影响因子: --
作者:
Rani, T. Sobha;Bapi, Raju S.
通讯作者: Bapi, Raju S.
DOI: 10.1186/1471-2105-11-s6-s17
发表时间: 2010-10-07
期刊: BMC bioinformatics
影响因子: 3
作者:
Bland C;Newsome AS;Markovets AA
通讯作者: Markovets AA
DOI: 10.1093/nar/gkl1024
发表时间: 2007
影响因子: 14.9
作者:
Mann S;Li J;Chen YP
通讯作者: Chen YP