WORDUP - AN EFFICIENT ALGORITHM FOR DISCOVERING STATISTICALLY SIGNIFICANT PATTERNS IN DNA-SEQUENCES

WORDUP - AN EFFICIENT ALGORITHM FOR DISCOVERING STATISTICALLY SIGNIFICANT PATTERNS IN DNA-SEQUENCES
复制标题

DOI:
10.1093/nar/20.11.2871
复制
发表时间:
1992-06-11
影响因子:
14.9
通讯作者:
SACCONE, C
SACCONE, C
中科院分区:
生物学2区
文献类型:
--
作者:
PESOLE, G;PRUNELLA, N;SACCONE, C

文献摘要

被引文献

相似文献

我们提出了一种快速灵敏的方法来分离具有非随机统计特性的短核苷酸序列,从而可能具有生物活性。它基于一阶马尔可夫分析,允许我们检测从6到10个核苷酸长的统计显着序列基序,这些基序在调查序列中显着共享(或避免)。该方法已在从真核生物启动子数据库中提取的521个序列上进行了测试(2)。我们的结果证明了该方法的准确性和效率,因为已知作为真核生物启动子的序列基序,如TATA-box和CAAT-box,被清楚地识别出来。此外,我们还发现了其他具有统计学意义的基序,其生物学作用尚待阐明。
We present here a fast and sensitive method designed to isolate short nucleotide sequences which have non-random statistical properties and may thus be biologically active. It is based on a first order Markov analysis and allows us to detect statistically significant sequence motifs from six to ten nucleotides long which are significantly shared (or avoided) in the sequences under investigation. This method has been tested on a set of 521 sequences extracted from the Eukaryotic Promoter Database (2). Our results demonstrate the accuracy and the efficiency of the method in that the sequence motifs which are known to act as eukaryotic promoters, such as the TATA-box and the CAAT-box, were clearly identified. In addition we have found other statistically significant motifs, the biological roles of which are yet to be clarified.