Adaptive Dictionary-based Compression of Protein Sequences

Adaptive Dictionary-based Compression of Protein Sequences
复制标题

基于自适应字典的蛋白质序列压缩

DOI:
10.5815/ijeme.2017.05.01
复制
发表时间:
2017
期刊:
International Journal of Education and Management Engineering
影响因子:
--
通讯作者:
Sunil Karforma
Sunil Karforma
中科院分区:
--
文献类型:
--
作者:
Akash Nag;Sunil Karforma

文献摘要

被引文献

相似文献

介绍了一种简单快速的蛋白质序列无损压缩算法--CAD。该算法特别适用于蛋白质组的压缩,蛋白质组是生物体表达的所有蛋白质的集合。该算法保持了活跃使用的氨基酸残基的不断变化的字典,结合自适应字典和霍夫曼编码,实现了每符号3.25比特的平均压缩比,比我们所知的大多数现有的蛋白质压缩和通用压缩算法要好。在平均压缩比为2.46:1,平均压缩比为1.32M残差/秒的情况下,该算法在压缩时间和压缩比上均优于其他任何一种蛋白质序列压缩算法。
This paper introduces a simple and fast lossless compression algorithm, called CAD, for the compression of protein sequences. The proposed algorithm is specially suited for compressing proteomes, which are the collection of all proteins expressed by an organism. Maintaining a changing dictionary of actively used amino-acid residues, the algorithm uses the adaptive dictionary together with Huffman coding to achieve an average compression rate of 3.25 bits per symbol, better than most other existing protein-compression and general-purpose compression algorithms known to us. With an average compression ratio of 2.46:1 and an average compression rate of 1.32M residues/sec, our algorithm outperforms every other compression algorithm for compressing protein sequences in terms of the balance in compression-time and compression rate.