Locality-based transfer learning on compression autoencoder for efficient scientific data lossy compression

Locality-based transfer learning on compression autoencoder for efficient scientific data lossy compression
复制标题

DOI:
10.1016/j.jnca.2022.103452
复制
发表时间:
2022-06
期刊:
J. Netw. Comput. Appl.
影响因子:
--
通讯作者:
Nan Wang;Tong Liu;Jinzhen Wang;Qing Liu;Shakeel Alibhai;Xubin He
Nan Wang;Tong Liu;Jinzhen Wang;Qing Liu;Shakeel Alibhai;Xubin He
中科院分区:
其他
文献类型:
--
作者:
Nan Wang;Tong Liu;Jinzhen Wang;Qing Liu;Shakeel Alibhai;Xubin He

文献摘要

相似文献

如今,科学模拟每次运行都能产生PB级的数据。为了显著减小数据大小,同时根据特定的用户要求保持压缩质量,SZ和ZFP等差错有界有损压缩技术现在变得流行起来。然而,这些技术仍然不能在低压缩误差的情况下实现超过两个数量级的压缩比。另一方面,在深度学习中,自动编码技术已被广泛应用于数据压缩,特别是图像压缩。作为一种替代方案,压缩自动编码器(CAE)最近被研究用于压缩科学数据。虽然CAE提供了比SZ和ZFP更高的压缩比,但它的训练开销很高,这使得它在实际压缩场景中几乎不实用。为了在获得较高压缩比的同时显著提高CAE的训练速度,本文提出了一种新的基于局部性的转移学习方法。我们还采用增量学习来保持较高的预测精度,并使用KL-散度作为指标来快速识别目标领域是否具有较低的测试误差。我们的评估结果表明,使用基于局部性的转移学习后,训练时间最多可以减少1200倍,并且仍然比目前最先进的科学数据有损压缩算法SZ的压缩比提高2到4倍。
Scientific simulation can generate petabyte-level data per run nowadays. To significantly reduce the data size while simultaneously maintaining the compression quality based on certain user requirements, error-bounded lossy compression techniques such as SZ and ZFP are now becoming popular. However, these techniques still cannot achieve a reduction ratio of more than two orders of magnitude with a low compression error. On the other hand, in deep learning, the autoencoder techniques have been widely used in data compression, especially images. As an alternative, the compression autoencoder (CAE) has recently been investigated to compress the scientific data. Although CAE provides a higher compression ratio than SZ and ZFP, it suffers from a high training overhead, which makes it almost impractical in real compression scenarios. In this paper, we propose a new locality-based transfer learning method in order to significantly increase the training speed of CAE while achieving a high compression ratio. We also adopt incremental learning to maintain a high prediction accuracy and use KL-divergence as an indicator to quickly identify whether a target domain has a low testing error. Our evaluation results show that, after using the locality-based transfer learning, the training time can be reduced by up to 1200 times, and still has a 2 to 4X compression ratio gain over the state-of-the-art scientific data lossy compressor SZ.