Optimizing Error-Bounded Lossy Compression for Scientific Data With Diverse Constraints

Optimizing Error-Bounded Lossy Compression for Scientific Data With Diverse Constraints
复制标题

优化具有不同约束的科学数据的误差有限有损压缩

DOI:
10.1109/tpds.2022.3194695
复制
发表时间:
2022
影响因子:
5.3
通讯作者:
Cappello, Franck
Cappello, Franck
中科院分区:
计算机科学2区
文献类型:
--
作者:
Liu, Yuanjian;Di, Sheng;Zhao, Kai;Jin, Sian;Wang, Cheng;Chard, Kyle;Tao, Dingwen;Foster, Ian;Cappello, Franck

文献摘要

参考文献

被引文献

相似文献

今天的科学模拟和先进仪器产生了大量的数据。由于I/O带宽、网络速度和存储容量有限,这些数据无法有效地存储和传输。误差受限有损压缩可以是解决这些问题的有效方法:它不仅可以显着减少数据大小,而且还可以根据用户定义的误差范围控制数据失真。在实践中,许多科学应用对有损压缩有特定的要求或约束,以保证重建的数据对于事后分析是有效的。例如,一些数据集包含不相关的数据,这些数据应该特别隔离,用户通常对值范围,地理空间区域和其他数据子集有直觉,这些数据子集对后续分析至关重要。然而,现有的最先进的误差受限有损压缩器在压缩期间不考虑这些约束,导致相对于用户的事后分析的较差压缩比,这是由于数据本身提供很少或没有事后分析价值的事实。在这项工作中,我们解决这个问题,提出了一个优化的框架,可以保持不同的限制,在错误有界的有损压缩,例如,清除不相关的数据,有效地保留多个值区间的不同精度,并允许用户在规则和不规则区域上设置不同的精度。我们在一台拥有多达2,100个内核的超级计算机上进行评估。六个实际应用的实验表明,我们提出的基于不同约束的误差有界有损压缩器可以获得更高的视觉质量或数据保真度的重建数据具有相同的或甚至更高的压缩比相比,传统的国家的最先进的压缩机SZ。我们的实验还表明,非常好的可扩展性的压缩性能相比,并行文件系统的I/O吞吐量。
Vast volumes of data are produced by today's scientific simulations and advanced instruments. These data cannot be stored and transferred efficiently because of limited I/O bandwidth, network speed, and storage capacity. Error-bounded lossy compression can be an effective method for addressing these issues: not only can it significantly reduce data size, but it can also control the data distortion based on user-defined error bounds. In practice, many scientific applications have specific requirements or constraints for lossy compression, in order to guarantee that the reconstructed data are valid for post hoc analysis. For example, some datasets contain irrelevant data that should be isolated in particular and users often have intuition regarding value ranges, geospatial regions, and other data subsets that are crucial for subsequent analysis. Existing state-of-the-art error-bounded lossy compressors, however, do not consider these constraints during compression, resulting in inferior compression ratios with respect to user's post hoc analysis, due to the fact that the data itself provides little or no value for post hoc analysis. In this work we address this issue by proposing an optimized framework that can preserve diverse constraints during the error-bounded lossy compression, e.g., cleaning the irrelevant data, efficiently preserving different precision for multiple value intervals, and allowing users to set diverse precision over both regular and irregular regions. We perform our evaluation on a supercomputer with up to 2,100 cores. Experiments with six real-world applications show that our proposed diverse constraints based error-bounded lossy compressor can obtain a higher visual quality or data fidelity on reconstructed data with the same or even higher compression ratios compared with the traditional state-of-the-art compressor SZ. Our experiments also demonstrate very good scalability in compression performance compared with the I/O throughput of the parallel file system.
DOI: 10.5194/gmd-12-4099-2019
发表时间: 2019-09-23
影响因子: 5.1
作者:
Delaunay, Xavier;Courtois, Aurelie;Gouillon, Flavien
通讯作者: Gouillon, Flavien
DOI: --
发表时间: 2017
期刊: 2017 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者:
Dingwen Tao;S. Di;Zizhong Chen;F. Cappello
通讯作者: F. Cappello
米兰达
DOI: --
发表时间: 2020
期刊: Cahiers Élisabéthains: A Journal of English Renaissance Studies
影响因子: --
作者:
Tina Krontiris
通讯作者: Tina Krontiris
DOI: --
发表时间: 2018
影响因子: 3.1
作者:
James Diffenderfer;Alyson Fox;J. Hittinger;G. Sanders;Peter Lindstrom
通讯作者: Peter Lindstrom
通过有损检查改善迭代方法的性能
DOI: --
发表时间: 2018
期刊: IEEE International Symposium on High-Performance Parallel Distributed Computing
影响因子: --
作者:
Dingwen Tao;S. Di;Xin Liang;Zizhong Chen;F. Cappello
通讯作者: F. Cappello