Human Pol II promoter recognition based on primary sequences and free energy of dinucleotides.

Human Pol II promoter recognition based on primary sequences and free energy of dinucleotides.
复制标题

基于一级序列和二核苷酸自由能的人 Pol II 启动子识别

DOI:
10.1186/1471-2105-9-113
复制
发表时间:
2008-02-24
期刊:
影响因子:
3
通讯作者:
Zhou LQ
Zhou LQ
中科院分区:
生物学4区
文献类型:
--
作者:
Yang JY;Zhou Y;Yu ZG;Anh V;Zhou LQ

文献摘要

参考文献

被引文献

相似文献

背景启动子区在决定特定基因转录起始位置方面起着重要作用。真核生物Pol II启动子序列的计算机预测是序列分析中最重要的问题之一。现有的启动子预测方法还远远不能令人满意。结果我们试图从由外显子和内含子组成的非启动子序列中识别出人类Pol II启动子序列。使用了四种方法:对由二核苷酸自由能、Z曲线分析和启动子/非启动子一级序列的全局描述子得到的数值序列进行了两种多重分形分析。从这些方法中提取了总共141个参数,并将其分为七组(方法)。它们用于产生某些空间,然后每个启动子/非启动子序列由相应空间中的点表示。测试了七种方法的所有120种可能的组合。基于Fisher线性判别算法,我们用相对较少的参数(96和117),得到了令人满意的判别精度。特别是在117个参数的情况下,训练集和测试集的准确率分别达到90.43%和89.79%。与其他五种现有方法的比较表明,我们的方法有更好的性能。使用全局描述符的方法(36个参数),17 18个实验验证的启动子序列的人类22号染色体被正确identified.ConclusionThe高精度实现表明,本文的方法是有用的理解困难的问题,启动子预测。
BackgroundPromoter region plays an important role in determining where the transcription of a particular gene should be initiated. Computational prediction of eukaryotic Pol II promoter sequences is one of the most significant problems in sequence analysis. Existing promoter prediction methods are still far from being satisfactory.ResultsWe attempt to recognize the human Pol II promoter sequences from the non-promoter sequences which are made up of exon and intron sequences. Four methods are used: two kinds of multifractal analysis performed on the numeric sequences obtained from the dinucleotide free energy, Z curve analysis and global descriptor of the promoter/non-promoter primary sequences. A total of 141 parameters are extracted from these methods and categorized into seven groups (methods). They are used to generate certain spaces and then each promoter/non-promoter sequence is represented by a point in the corresponding space. All the 120 possible combinations of the seven methods are tested. Based on Fisher's linear discriminant algorithm, with a relatively smaller number of parameters (96 and 117), we get satisfactory discriminant accuracies. Particularly, in the case of 117 parameters, the accuracies for the training and test sets reach 90.43% and 89.79%, respectively. A comparison with five other existing methods indicates that our methods have a better performance. Using the global descriptor method (36 parameters), 17 of the 18 experimentally verified promoter sequences of human chromosome 22 are correctly identified.ConclusionThe high accuracies achieved suggest that the methods of this paper are useful for understanding the difficult problem of promoter prediction.
DOI: 10.1186/1471-2105-7-9
发表时间: 2006-01-10
期刊: BMC bioinformatics
影响因子: 3
作者:
Guo FB;Zhang CT
通讯作者: Zhang CT
DOI: 10.1073/pnas.92.19.8700
发表时间: 1995-09-12
影响因子: 11.1
作者:
DUBCHAK, I;MUCHNIK, I;KIM, SH
通讯作者: KIM, SH
DOI: 10.1093/nar/gki271
发表时间: 2005-01-01
影响因子: 14.9
作者:
Gangal, R;Sharma, P
通讯作者: Sharma, P
DOI: 10.1093/nar/gkg254
发表时间: 2003-03-15
影响因子: 14.9
作者:
Guo, FB;Ou, HY;Zhang, CT
通讯作者: Zhang, CT
DOI: 10.1101/gr.10.4.539
发表时间: 2000-04-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Ohler, U
通讯作者: Ohler, U