Adaptive Dictionary-based Compression of Protein Sequences
Adaptive Dictionary-based Compression of Protein Sequences
复制标题
基于自适应字典的蛋白质序列压缩
DOI:
10.5815/ijeme.2017.05.01
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
Sunil Karforma
中科院分区:
文献类型:
--
作者:
Akash Nag;Sunil Karforma
This paper introduces a simple and fast lossless compression algorithm, called CAD, for the compression of protein sequences. The proposed algorithm is specially suited for compressing proteomes, which are the collection of all proteins expressed by an organism. Maintaining a changing dictionary of actively used amino-acid residues, the algorithm uses the adaptive dictionary together with Huffman coding to achieve an average compression rate of 3.25 bits per symbol, better than most other existing protein-compression and general-purpose compression algorithms known to us. With an average compression ratio of 2.46:1 and an average compression rate of 1.32M residues/sec, our algorithm outperforms every other compression algorithm for compressing protein sequences in terms of the balance in compression-time and compression rate.