Error-Controlled, Progressive, and Adaptable Retrieval of Scientific Data with Multilevel Decomposition

Error-Controlled, Progressive, and Adaptable Retrieval of Scientific Data with Multilevel Decomposition
复制标题

通过多级分解进行误差控制、渐进且适应性强的科学数据检索

DOI:
10.1145/3458817.3476179
复制
发表时间:
2021
期刊:
SC21: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
S. Klasky
S. Klasky
中科院分区:
--
文献类型:
--
作者:
Xin Liang;Qian Gong;Jieyang Chen;Ben Whitney;Lipeng Wan;Qing Liu;D. Pugmire;Rick Archibald;N. Podhorszki;S. Klasky

文献摘要

参考文献

被引文献

相似文献

极端规模的模拟和高分辨率仪器产生了越来越多的数据,这不仅对运行期间的数据存储提出了重大挑战,而且对数据将在很长一段时间内重复检索和分析的后处理也提出了重大挑战。在满足广泛的事后分析需求的同时,最大限度地减少由不适当和/或过度的数据检索引起的I/O开销,这一挑战永远不应被忽视。在本文中,我们提出了一个数据重构,压缩和检索框架,能够1)细粒度的数据重构精度方面; 2)增量检索和重组的数据在各种误差范围;和3)自适应检索数据在多精度和多分辨率的不同分析。随着渐进的数据重组和自适应的检索算法,我们的框架显着减少了检索的数据量时,多个增量精度的要求和/或下游的分析时间时,使用粗分辨率。实验表明,在相同的渐进式请求的错误界限下,使用我们的框架检索的数据量比使用最先进的单一错误界限的方法少64%。在1024个核和$\sim\ 600$ GB数据的并行实验中,我们的方法在写入和阅读持久存储系统时分别比现有方法产生了1.36\times $和2.52\times $的性能。
Extreme-scale simulations and high-resolution instruments have been generating an increasing amount of data, which poses significant challenges to not only data storage during the run, but also post-processing where data will be repeatedly retrieved and analyzed for a long period of time. The challenges in satisfying a wide range of post-hoc analysis needs while minimizing the I/O overhead caused by inappropriate and/or excessive data retrieval should never be left unmanaged. In this paper, we propose a data refactoring, compressing, and retrieval framework capable of 1) fine-grained data refactoring with regard to precision; 2) incrementally retrieving and recomposing the data in terms of various error bounds; and 3) adaptively retrieving data in multi-precision and multi-resolution with respect to different analysis. With the progressive data re-composition and the adaptable retrieval algorithms, our framework significantly reduces the amount of data retrieved when multiple incremental precision are requested and/or the downstream analysis time when coarse resolution is used. Experiments show that the amount of data retrieved under the same progressively requested error bound using our framework is 64% less than that using state-of-the-art single-error-bounded approaches. Parallel experiments with up to 1, 024 cores and $\sim\ 600$ GB data in total show that our approach yields $1.36\times$ and $2.52\times$ performance over existing approaches in writing to and reading from persistent storage systems, respectively.
AMM:自适应多线性网格
DOI: 10.1109/tvcg.2022.3165392
发表时间: 2022
影响因子: 5.2
作者:
Bhatia, Harsh;Hoang, Duong;Morrical, Nate;Pascucci, Valerio;Bremer, Peer-Timo;Lindstrom, Peter
通讯作者: Lindstrom, Peter