High-Ratio Lossy Compression: Exploring the Autoencoder to Compress Scientific Data

High-Ratio Lossy Compression: Exploring the Autoencoder to Compress Scientific Data
复制标题

DOI:
10.1109/tbdata.2021.3066151
复制
发表时间:
2023-02
影响因子:
7.2
通讯作者:
Tong Liu;Jinzhen Wang;Qing Liu;Shakeel Alibhai;Tao Lu;Xubin He
Tong Liu;Jinzhen Wang;Qing Liu;Shakeel Alibhai;Tao Lu;Xubin He
中科院分区:
计算机科学2区
文献类型:
--
作者:
Tong Liu;Jinzhen Wang;Qing Liu;Shakeel Alibhai;Tao Lu;Xubin He

文献摘要

被引文献

相似文献

高性能计算(HPC)系统上的科学模拟可以每次运行生成大量的浮点数据。为了减轻数据存储瓶颈并降低数据量,使用浮点数压缩机是常见的。与无损压缩机相比,有损压缩机(例如SZ和ZFP)可以在保持数据的实用性的同时更加积极地减少数据量。但是,如果不严重扭曲数据,几乎不可能将超过两个数量级的降低比率降低。在深度学习中,自动编码器技术表现出很大的数据压缩潜力,尤其是图像。但是,自动编码器是否可以在科学数据上提供类似的性能。在本文中,我们首次对使用自动编码器来压缩现实世界的科学数据并说明了使用自动编码器减少科学数据的几个关键发现的全面研究。我们实现了基于自动编码器的压缩原型,以减少浮点数据。我们的研究表明,需要进一步调整开箱即用的实现,以达到高压比和令人满意的误差范围。我们的评估结果表明,对于大多数测试数据集,调整的自动编码器的表现分别高于4倍,而ZFP的压缩比分别高达50倍。我们在这项工作中学到的实践和经验教训可以指导未来使用自动编码器压缩科学数据的优化。
Scientific simulations on high-performance computing (HPC) systems can generate large amounts of floating-point data per run. To mitigate the data storage bottleneck and lower the data volume, it is common for floating-point compressors to be employed. As compared to lossless compressors, lossy compressors, such as SZ and ZFP, can reduce data volume more aggressively while maintaining the usefulness of the data. However, a reduction ratio of more than two orders of magnitude is almost impossible without seriously distorting the data. In deep learning, the autoencoder technique has shown great potential for data compression, in particular with images. Whether the autoencoder can deliver similar performance on scientific data, however, is unknown. In this article, we for the first time conduct a comprehensive study on the use of autoencoders to compress real-world scientific data and illustrate several key findings on using autoencoders for scientific data reduction. We implement an autoencoder-based compression prototype to reduce floating-point data. Our study shows that the out-of-the-box implementation needs to be further tuned in order to achieve high compression ratios and satisfactory error bounds. Our evaluation results show that, for most of the test datasets, the tuned autoencoder outperforms SZ by up to 4X, and ZFP by up to 50X in compression ratios, respectively. Our practices and lessons learned in this work can direct future optimizations for using autoencoders to compress scientific data.