SPDP: An Automatically Synthesized Lossless Compression Algorithm for Floating-Point Data

SPDP: An Automatically Synthesized Lossless Compression Algorithm for Floating-Point Data
复制标题

SPDP:一种自动合成的浮点数据无损压缩算法

DOI:
10.1109/dcc.2018.00042
复制
发表时间:
2018
期刊:
2018 Data Compression Conference
影响因子:
--
通讯作者:
Martin Burtscher
Martin Burtscher
中科院分区:
--
文献类型:
--
作者:
S. Claggett;S. Azimi;Martin Burtscher

文献摘要

被引文献

相似文献

科学计算产生、传输和存储大量单精度和双精度浮点数据,使其成为可以从数据压缩中受益匪浅的领域。为了深入了解如何为此类数据提供有效的无损压缩算法,我们生成了超过900万种算法,并选择了在26个数据集上产生最高压缩比的算法。由此产生的算法,称为SPDP,包括四个数据转换,专门在字或字节粒度操作。尽管如此,SPDP在11个数据集上提供了最高的压缩比,并且平均而言,除了七个比较的压缩器中的一个之外,其他所有压缩器都优于SPDP。对SPDP内部的分析揭示了如何为科学数据构建有效的压缩算法。
Scientific computing produces, transfers, and stores massive amounts of single- and double-precision floating-point data, making this a domain that can greatly benefit from data compression. To gain insight into what makes an effective lossless compression algorithm for such data, we generated over nine million algorithms and selected the one that yields the highest compression ratio on 26 datasets. The resulting algorithm, called SPDP, comprises four data transformations that operate exclusively at word or byte granularity. Nevertheless, SPDP delivers the highest compression ratio on eleven datasets and, on average, outperforms all but one of the seven compared compressors. An analysis of SPDP's internals reveals how to build effective compression algorithms for scientific data.