STATISTICAL PROPERTIES OF DNA-SEQUENCES

STATISTICAL PROPERTIES OF DNA-SEQUENCES
复制标题

DOI:
10.1016/0378-4371(95)00247-5
复制
发表时间:
1995-11-15
期刊:
影响因子:
3.3
通讯作者:
STANLEY, HE
STANLEY, HE
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
PENG, CK;BULDYREV, SV;STANLEY, HE

文献摘要

被引文献

相似文献

我们回顾证据支持的想法,在基因中含有非编码区的DNA序列是相关的,并且相关性是非常长的范围-事实上,核苷酸数千碱基对远相关。我们在基因编码区没有发现这样的长程相关性。我们通过应用一种新的算法--去趋势波动分析(DFA)来解决碱基对序列的“非平稳性”特征问题。我们解决的说法,有没有差异的DNA编码区和非编码区的统计特性,系统地应用DFA算法,以及标准的FFT分析,在整个GenBank数据库中的每个DNA序列(33 301编码和29 453非编码)。最后,我们简要介绍了最近的一些工作表明,非编码序列有一定的统计特征,在自然和人工语言的共同点。具体来说,我们适应DNA的Zipf方法来分析语言文本,这些统计性质的非编码序列支持的可能性,DNA的非编码区可能携带生物信息。
We review evidence supporting the idea that the DNA sequence in genes containing non-coding regions is correlated, and that the correlation is remarkably long range - indeed, nucleotides thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene, We resolve the problem of the ''non-stationarity'' feature of the sequence of base pairs by applying a new algorithm called detrended fluctuation analysis (DFA). We address the claim of Voss that there is no difference in the statistical properties of coding and non-coding regions of DNA by systematically applying the DFA algorithm, as well as standard FFT analysis, to every DNA sequence (33 301 coding and 29 453 non-coding) in the entire GenBank database. Finally, we describe briefly some recent work showing that the noncoding sequences have certain statistical features in common with natural and artificial languages. Specifically, we adapt to DNA the Zipf approach to analyzing linguistic texts, These statistical properties of non-coding sequences support the possibility that non-coding regions of DNA may carry biological information.