Signi cantly Lower Entropy Estimates for Natural DNA Sequences

Signi cantly Lower Entropy Estimates for Natural DNA Sequences
复制标题

显着降低自然 DNA 序列的熵估计

DOI:
--
复制
发表时间:
1996
期刊:
影响因子:
--
通讯作者:
P. Yianilos
P. Yianilos
中科院分区:
--
文献类型:
--
作者:
David Loewenstern;P. Yianilos

文献摘要

被引文献

相似文献

如果DNA是在其字母表fA;C;G;Tg上的随机字符串,则最佳代码将为每个核苷酸分配2位。我们把DNA想象成一种高度有序的、有目的的分子,因此可以合理地期望它的弦表示的统计模型产生低得多的熵估计。令人惊讶的是,许多天然DNA序列,包括人类基因组的一部分,情况并非如此。我们介绍了一种新的统计模型(压缩算法),迄今为止报道的最强的,自然发生的DNA序列。传统的技术编码一个核苷酸使用的比特数(1.90)比仅依靠单个核苷酸的频率统计(1.95)获得的比特数略少。我们的方法在某些情况下增加了超过五倍(1.66)的差距,并可能导致更好的性能在微生物模式识别应用。我们的主要贡献之一,以及这些改进的主要来源,是在模型中正式包含不精确的匹配信息。在不同距离上存在的匹配形成了一个专家小组,然后将其组合成一个单一的预测。这种组合的结构是新颖的,它的参数是使用期望最大化(EM)学习。
If DNA were a random string over its alphabet fA;C;G;Tg, an optimal code would assign 2 bits to each nucleotide. We imagine DNA to be a highly ordered, purposeful molecule, and might therefore reasonably expect statistical models of its string representation to produce much lower entropy estimates. Surprisingly this has not been the case for many natural DNA sequences, including portions of the human genome. We introduce a new statistical model (compression algorithm), the strongest reported to date, for naturally occurring DNA sequences. Conventional techniques code a nucleotide using only slightly fewer bits (1.90) than one obtains by relying only on the frequency statistics of individual nucleotides (1.95). Our method in some cases increases this gap by more than ve-fold (1.66) and may lead to better performance in microbiological pattern recognition applications. One of our main contributions, and the principle source of these improvements, is the formal inclusion of inexact match information in the model. The existence of matches at various distances forms a panel of experts which are then combined into a single prediction. The structure of this combination is novel and its parameters are learned using Expectation Maximization (EM).