System-Level Scalable Checkpoint-Restart for Petascale Computing

System-Level Scalable Checkpoint-Restart for Petascale Computing
复制标题

用于千万亿次计算的系统级可扩展检查点重启

DOI:
10.1109/icpads.2016.0125
复制
发表时间:
2016
期刊:
2016 IEEE 22nd International Conference on Parallel and Distributed Systems (ICPADS)
影响因子:
--
通讯作者:
G. Cooperman
G. Cooperman
中科院分区:
--
文献类型:
--
作者:
Jiajun Cao;K. Arya;Rohan Garg;Shawn Matott;D. Panda;H. Subramoni;Jérôme Vienne;G. Cooperman

文献摘要

被引文献

相似文献

对于即将到来的艾级世代的容错一直是一个活跃的研究领域。容错策略的组成部分之一是检查点。Petascale级别的检查点是通过一种新的机制来演示的,该机制用于虚拟化InfiniBand UD(不可靠数据报)模式,并用于在每个基于UD的发送上更新远程地址,因为缺乏固定的对等体。请注意,需要InfiniBand UD来支持现代MPI实现。从目前的结果到未来的SSD存储系统的外推提供了证据,目前的方法将保持实用的艾级一代。这种透明的检查点方法使用DMTCP检查点包的框架进行评估。结果显示为HPCG(线性代数),NAMD(分子动力学),和NAS NPB基准。在32,752个CPU内核上测试了多达32,752个MPI进程,演示了在11分钟内使用38 TB内存占用进行计算的检查点操作。将系统开销降低到1%以下。该方法还评估了三个广泛使用的MPI实现。
Fault tolerance for the upcoming exascale generation has long been an area of active research. One of the components of a fault tolerance strategy is checkpointing. Petascale-level checkpointing is demonstrated through a new mechanism for virtualization of the InfiniBand UD (unreliable datagram) mode, and for updating the remote address on each UD-based send, due to lack of a fixed peer. Note that InfiniBand UD is required to support modern MPI implementations. An extrapolation from the current results to future SSD-based storage systems provides evidence that the current approach will remain practical in the exascale generation. This transparent checkpointing approach is evaluated using a framework of the DMTCP checkpointing package. Results are shown for HPCG (linear algebra), NAMD (molecular dynamics), and the NAS NPB benchmarks. In tests up to 32,752 MPI processes on 32,752 CPU cores, checkpointing of a computation with a 38 TB memory footprint in 11 minutes is demonstrated. Runtime overhead is reduced to less than 1%. The approach is also evaluated across three widely used MPI implementations.