Protein is incompressible

Protein is incompressible
复制标题

蛋白质是不可压缩的

DOI:
10.1109/dcc.1999.755675
复制
发表时间:
1999
期刊:
Proceedings DCC'99 Data Compression Conference (Cat. No. PR00096)
影响因子:
--
通讯作者:
I. Witten
I. Witten
中科院分区:
--
文献类型:
--
作者:
C. Nevill;I. Witten

文献摘要

被引文献

相似文献

生命基于两种聚合物,DNA和蛋白质,它们的性质可以用一个简单的文本文件来描述。标准文本压缩技术对生物序列的作用与对英文文本的作用一样,这是很自然的。但生物序列具有与语言序列根本不同的结构,标准的压缩方案在它们上表现出令人失望的性能。我们描述了一种新的压缩方法,它考虑了潜在的生化原理。这导致了统计压缩器的混合的一般化,其中使用每个上下文,通过其与当前上下文的相似性来加权。结果支持了生物信息学的研究表明,蛋白质中几乎没有马尔可夫依赖性。这削弱了数据压缩方案,并将它们减少到零阶模型。
Life is based on two polymers, DNA and protein, whose properties can be described in a simple text file. It is natural to expect that standard text compression techniques would work on biological sequences as they do on English text. But biological sequences have a fundamentally different structure from linguistic ones, and standard compression schemes exhibit disappointing performance on them. We describe a new approach to compression that takes account of the underlying biochemical principles. This gives rise to a generalization of blending for statistical compressors where every context is used, weighted by its similarity to the current context. Results support what research in bioinformatics has shown, that there is little Markov dependency in protein. This cripples data compression schemes and reduces them to order zero models.