Toward Quantity-of-Interest Preserving Lossy Compression for Scientific Data

Toward Quantity-of-Interest Preserving Lossy Compression for Scientific Data
复制标题

DOI:
10.14778/3574245.3574255
复制
发表时间:
2022-12
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Pu Jiao;S. Di;Hanqi Guo;Kai Zhao;Jiannan Tian;Dingwen Tao;Xin Liang;F. Cappello
Pu Jiao;S. Di;Hanqi Guo;Kai Zhao;Jiannan Tian;Dingwen Tao;Xin Liang;F. Cappello
中科院分区:
其他
文献类型:
--
作者:
Pu Jiao;S. Di;Hanqi Guo;Kai Zhao;Jiannan Tian;Dingwen Tao;Xin Liang;F. Cappello

文献摘要

被引文献

相似文献

今天的科学模拟和仪器正在产生大量的数据,导致存储、传输和分析这些数据的困难。虽然差错控制的有损压缩器在显著减少数据量和高效开发用于多种科学应用的数据库方面是有效的,但它们主要支持对原始数据的差错控制,这在数据和用户的下游分析之间留下了很大的差距。这可能会在分析结果中造成不合格的不确定性,也就是兴趣量(QOI),这是用户在实践中采用有损压缩时的主要担忧。在这篇文章中,我们提出了严格的数学理论来保存四类在有损压缩过程中广泛应用于科学分析的QOI,并给出了具体的实现方法。具体地说,我们首先发展了单变量QOI的误差控制理论,这是计算动能等物理性质所必需的,然后是在现实世界应用中更常用的多变量QOI。该方法以模块化的方式集成到最先进的压缩框架中,可以很容易地适应新的QOI和新的压缩算法。在真实数据集上的实验表明,所提出的方法在没有试验和错误的情况下,对包括动能、区域平均和等值面在内的重要的QOI提供了忠实的错误控制,同时提供了高达最新压缩比的4倍的压缩比。
Today's scientific simulations and instruments are producing a large amount of data, leading to difficulties in storing, transmitting, and analyzing these data. While error-controlled lossy compressors are effective in significantly reducing data volumes and efficiently developing databases for multiple scientific applications, they mainly support error controls on raw data, which leaves a significant gap between the data and user's downstream analysis. This may cause unqualified uncertainties in the outcomes of the analysis, a.k.a quantities of interest (QoIs), which are the major concerns of users in adopting lossy compression in practice. In this paper, we propose rigorous mathematical theories to preserve four families of QoIs that are widely used in scientific analysis during lossy compression along with practical implementations. Specifically, we first develop the error control theory for univariate QoIs which are essential for computing physical properties such as kinetic energy, followed by multivariate QoIs that are more commonly used in real-world applications. The proposed method is integrated into a state-of-the-art compression framework in a modular fashion, which could easily adapt to new QoIs and new compression algorithms. Experiments on real-world datasets demonstrate that the proposed method provides faithful error control on important QoIs including kinetic energy, regional average, and isosurface without trials and errors, while offering compression ratios that are up to 4X of the compression ratios provided by state-of-the-art compressors.