A fractal method to distinguish coding and non-coding sequences in a complete genome based on a number sequence representation

A fractal method to distinguish coding and non-coding sequences in a complete genome based on a number sequence representation
复制标题

DOI:
10.1016/j.jtbi.2004.09.002
复制
发表时间:
2005-02-21
影响因子:
2
通讯作者:
Long, SC
Long, SC
中科院分区:
生物学4区
文献类型:
--
作者:
Zhou, LQ;Yu, ZG;Long, SC

文献摘要

被引文献

相似文献

根据全基因组中编码序列和非编码序列统计特性的不同,提出了一种分形方法来区分这两类序列。我们首先提出了一个数字序列表示的DNA序列。然后对得到的数列的测度表示进行多重分形分析。从多重分形分析的结果中选取了三个指数C-1、C-1和C-2。每个DNA可以由这些三分量矢量生成的三维空间中的点表示。结果表明,许多原核生物全基因组中编码序列和非编码序列对应的点大致分布在不同的区域。Fisher判别算法可用于在跨越空间中分离这两个区域。如果DNA序列的点(C-1,C-1,C-2)位于对应于编码序列的区域中,则该序列被鉴别为编码序列;否则,该序列被分类为非编码序列。对所有51种原核生物,平均判别准确率p(c)、p(nc)、q(c)和q(nc)分别达到72.28%、84.65%、72.53%和84.18%。(C)2004 Elsevier Ltd.保留所有权利。
A fractal method to distinguish coding and non-coding sequences in a complete genome is proposed, based on different statistical behaviors between these two kinds of sequences. We first propose a number sequence representation of DNA sequences. Multifractal analysis is then performed on the measure representation of the obtained number sequence. The three exponents C-1, C-1 and C-2 are selected from the result of multifractal analysis. Each DNA may be represented by a point in the three-dimensional space generated by these three-component vectors. It is shown that points corresponding to coding and non-coding sequences in the complete genome of many prokaryotes are roughly distributed in different regions. Fisher's discriminant algorithm can be used to separate these two regions in the spanned space. If the point (C-1,C-1,C-2) for a DNA sequence is situated in the region corresponding to coding sequences, the sequence is discriminated as a coding sequence; otherwise, the sequence is classified as a non-coding one. For all 51 prokaryotes we considered, the average discriminant accuracies p(c), p(nc), q(c), and q(nc), reach 72.28%, 84.65%, 72.53% and 84.18%, respectively. (C) 2004 Elsevier Ltd. All rights reserved.