Understanding and Modeling Lossy Compression Schemes on HPC Scientific Data

Understanding and Modeling Lossy Compression Schemes on HPC Scientific Data
复制标题

DOI:
10.1109/ipdps.2018.00044
复制
发表时间:
2018-05
期刊:
2018 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Tao Lu;Qing Liu;Xubin He;Huizhang Luo;E. Suchyta;J. Choi;N. Podhorszki;S. Klasky;M. Wolf;Tong Liu;Zhenbo Qiao
Tao Lu;Qing Liu;Xubin He;Huizhang Luo;E. Suchyta;J. Choi;N. Podhorszki;S. Klasky;M. Wolf;Tong Liu;Zhenbo Qiao
中科院分区:
其他
文献类型:
--
作者:
Tao Lu;Qing Liu;Xubin He;Huizhang Luo;E. Suchyta;J. Choi;N. Podhorszki;S. Klasky;M. Wolf;Tong Liu;Zhenbo Qiao

文献摘要

被引文献

相似文献

科学模拟会产生大量的浮点数据,使用传统的精简方案(如去重或无损压缩)往往不太容易压缩这些数据。有损浮点压缩的出现有望满足高性能计算应用对数据精简的需求;然而,有损压缩在科学成果产出中尚未得到广泛应用。我们认为一个根本原因是对有损压缩在科学数据上的益处、缺陷以及性能缺乏了解。在本文中,我们使用真实且具有代表性的高性能计算数据集,对最先进的有损压缩技术(包括ZFP、SZ和ISABELA)进行了全面研究。我们的评估揭示了压缩器设计、数据特征和压缩性能之间复杂的相互作用。通过对融合斑点检测的案例研究,我们还检验了精度降低对数据分析的影响,为相关领域的科学家提供了关于保真度损失预期的见解。此外,通过反复试验来了解压缩性能会涉及大量的计算和存储开销。为此,我们提出了一种基于采样的估算方法,该方法从数据样本中推断出压缩率,以指导相关领域的科学家做出更明智的数据精简决策。
Scientific simulations generate large amounts of floating-point data, which are often not very compressible using the traditional reduction schemes, such as deduplication or lossless compression. The emergence of lossy floating-point compression holds promise to satisfy the data reduction demand from HPC applications; however, lossy compression has not been widely adopted in science production. We believe a fundamental reason is that there is a lack of understanding of the benefits, pitfalls, and performance of lossy compression on scientific data. In this paper, we conduct a comprehensive study on state-of-the-art lossy compression, including ZFP, SZ, and ISABELA, using real and representative HPC datasets. Our evaluation reveals the complex interplay between compressor design, data features and compression performance. The impact of reduced accuracy on data analytics is also examined through a case study of fusion blob detection, offering domain scientists with the insights of what to expect from fidelity loss. Furthermore, the trial and error approach to understanding compression performance involves substantial compute and storage overhead. To this end, we propose a sampling based estimation method that extrapolates the reduction ratio from data samples, to guide domain scientists to make more informed data reduction decisions.