Maximizing I/O Bandwidth for Reverse Time Migration on Heterogeneous Large-Scale Systems

Maximizing I/O Bandwidth for Reverse Time Migration on Heterogeneous Large-Scale Systems
复制标题

最大化异构大型系统上逆时迁移的 I/O 带宽

DOI:
10.1007/978-3-030-57675-2_17
复制
发表时间:
2020
期刊:
Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing
影响因子:
--
通讯作者:
D. Keyes
D. Keyes
中科院分区:
--
文献类型:
--
作者:
Tariq Alturkestani;H. Ltaief;D. Keyes

文献摘要

被引文献

相似文献

逆时偏移(RTM)是油气勘探的一项重要科学应用。3D RTM模拟生成的中间数据太多了,无法放入主内存。特别地,RTM具有两个连续的计算阶段,即,前向建模和后向传播,这需要在时间积分的特定时间步长写入然后读取计算的解网格的状态。存储器体系结构的进步使得在大规模系统上集成分层存储介质变得可行和可负担,从传统的并行文件系统(PFS)开始到中间快速磁盘技术(例如,节点本地和远程共享的突发缓冲区)和CPU主存储器。为了应对异构HPC系统部署的趋势,我们引入了对多层缓冲系统(MLBS)框架的扩展,以在GPU硬件加速器存在的情况下进一步最大化RTM I/O带宽。其主要思想是利用GPU的高带宽内存(HBM)作为额外的存储介质层。MLBS的目标最终是通过启用跨所有分层存储介质层操作的缓冲机制来隐藏应用程序的I/O开销。MLBS因此能够维持每个存储介质层的I/O带宽。通过异步地执行昂贵的I/O操作并为数据运动与计算重叠创造机会,MLBS可以将RTM应用的原始I/O绑定行为转换为计算绑定机制。事实上,MLBS的预取策略允许RTM应用程序相信它可以访问GPU上更大的内存容量,同时透明地跨存储层执行必要的内务处理。我们证明了MLBS的有效性首脑会议上的超级计算机使用2048个计算节点,配备了12288个GPU,通过实现高达1.4倍的性能加速相比,参考PFS为基础的RTM实现大型3D解决方案网格。
Reverse Time Migration (RTM) is an important scientific application for oil and gas exploration. The 3D RTM simulation generates terabytes of intermediate data that does not fit in main memory. In particular, RTM has two successive computational phases, i.e., the forward modeling and the backward propagation, that necessitate to write and then to read the state of the computed solution grid at specific time steps of the time integration. Advances in memory architecture have made it feasible and affordable to integrate hierarchical storage media on large-scale systems, starting from the traditional Parallel File Systems (PFS) to intermediate fast disk technologies (e.g., node-local and remote-shared Burst Buffer) and up to CPU main memory. To address the trend of heterogeneous HPC systems deployment, we introduce an extension to our Multilayer Buffer System (MLBS) framework to further maximize RTM I/O bandwidth in presence of GPU hardware accelerators. The main idea is to leverage the GPU’s High Bandwidth Memory (HBM) as an additional storage media layer. The objective of MLBS is ultimately to hide the application’s I/O overhead by enabling a buffering mechanism operating across all the hierarchical storage media layers. MLBS is therefore able to sustain the I/O bandwidth at each storage media layer. By asynchronously performing expensive I/O operations and creating opportunities for overlapping data motion with computations, MLBS may transform the original I/O bound behavior of the RTM application into a compute-bound regime. In fact, the prefetching strategy of MLBS allows the RTM application to believe that it has access to a larger memory capacity on the GPU, while transparently performing the necessary housekeeping across the storage layers. We demonstrate the effectiveness of MLBS on the Summit supercomputer using 2048 compute nodes equipped with a total of 12288 GPUs by achieving up to 1.4X performance speedup compared to the reference PFS-based RTM implementation for large 3D solution grid.