Region-adaptive, Error-controlled Scientific Data Compression using Multilevel Decomposition

Region-adaptive, Error-controlled Scientific Data Compression using Multilevel Decomposition
复制标题

DOI:
10.1145/3538712.3538717
复制
发表时间:
2022-07
期刊:
Proceedings of the 34th International Conference on Scientific and Statistical Database Management
影响因子:
--
通讯作者:
Qian Gong;Ben Whitney;Chengzhu Zhang;Xin Liang;A. Rangarajan;Jieyang Chen;Lipeng Wan;P. Ullrich;Qing Liu;R. Jacob;Sanjay Ranka;S. Klasky
Qian Gong;Ben Whitney;Chengzhu Zhang;Xin Liang;A. Rangarajan;Jieyang Chen;Lipeng Wan;P. Ullrich;Qing Liu;R. Jacob;Sanjay Ranka;S. Klasky
中科院分区:
其他
文献类型:
--
作者:
Qian Gong;Ben Whitney;Chengzhu Zhang;Xin Liang;A. Rangarajan;Jieyang Chen;Lipeng Wan;P. Ullrich;Qing Liu;R. Jacob;Sanjay Ranka;S. Klasky

文献摘要

被引文献

相似文献

计算机处理速度的提高大大超过了网络和存储带宽的提高,导致现代科学面临大数据挑战,科学应用可以快速生成比可以传输和存储的数据更多的数据。因此,大科学数据必须减少几个数量级,而减少的数据的准确性需要保证进一步的科学探索。此外,科学家们往往感兴趣的是一些特定的空间/时间区域的数据,在那里需要更高的精度。需要高精度的区域的位置有时可以基于应用知识来规定,而其他时候它们必须基于一般的空间/时间变化来估计。在本文中,我们开发了一种新的多层次的方法,允许用户施加区域明智的压缩误差范围。我们的方法利用一个多级压缩机的副产品,以检测区域的细节是丰富的,我们提供的理论基础,区域明智的错误控制。随着空间变化的精度保存,我们的方法可以实现显着更高的压缩比单误差有界压缩方法和控制错误的兴趣区域。我们对两个气候用例进行了评估-一个针对小规模的节点特征,另一个侧重于长期的区域特征。对于这两种用例,在压缩之前,特征的位置是未知的。通过基于多尺度空间变化选择大约16%的数据,并以比其余区域更小的误差容限压缩这些区域,与相同压缩比的单误差有界压缩相比,我们的方法将后分析的准确性提高了大约2倍。使用相同的误差范围的区域的兴趣,我们的方法可以实现超过50%的整体压缩比的增加。
The increase of computer processing speed is significantly outpacing improvements in network and storage bandwidth, leading to the big data challenge in modern science, where scientific applications can quickly generate much more data than that can be transferred and stored. As a result, big scientific data must be reduced by a few orders of magnitude while the accuracy of the reduced data needs to be guaranteed for further scientific explorations. Moreover, scientists are often interested in some specific spatial/temporal regions in their data, where higher accuracy is required. The locations of the regions requiring high accuracy can sometimes be prescribed based on application knowledge, while other times they must be estimated based on general spatial/temporal variation. In this paper, we develop a novel multilevel approach which allows users to impose region-wise compression error bounds. Our method utilizes the byproduct of a multilevel compressor to detect regions where details are rich and we provide the theoretical underpinning for region-wise error control. With spatially varying precision preservation, our approach can achieve significantly higher compression ratios than single-error bounded compression approaches and control errors in the regions of interest. We conduct the evaluations on two climate use cases – one targeting small-scale, node features and the other focusing on long, areal features. For both use cases, the locations of the features were unknown ahead of the compression. By selecting approximately 16% of the data based on multi-scale spatial variations and compressing those regions with smaller error tolerances than the rest, our approach improves the accuracy of post-analysis by approximately 2 × compared to single-error-bounded compression at the same compression ratio. Using the same error bound for the region of interest, our approach can achieve an increase of more than 50% in overall compression ratio.