Interconnect agnostic checkpoint/restart in open MPI
Interconnect agnostic checkpoint/restart in open MPI
复制标题
开放 MPI 中的互连不可知检查点/重启
DOI:
10.1145/1551609.1551619
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
A. Lumsdaine
中科院分区:
文献类型:
--
作者:
Joshua Hursey;T. Mattox;A. Lumsdaine
Long running High Performance Computing (HPC) applications at scale must be able to tolerate inevitable faults if they are to harness current and future HPC systems. Message Passing Interface (MPI) level transparent checkpoint/restart fault tolerance is an appealing option to HPC application developers that do not wish to restructure their code. Historically, MPI implementations that provided this option have struggled to provide a full range of interconnect support, especially shared memory support. This paper presents a new approach for implementing checkpoint/restart coordination algorithms that allows the MPI implementation of checkpoint/restart to be interconnect agnostic. This approach allows an application to be checkpointed on one set of interconnects (e.g., InfiniBand and shared memory) and be restarted with a different set of interconnects (e.g., Myrinet and shared memory or Ethernet). By separating the network interconnect details from the checkpoint/restart coordination algorithm we allow the HPC application to respond to changes in the cluster environment such as interconnect unavailability due to switch failure, re-load balance on an existing machine, or migrate to a different machine with a different set of interconnects. We present results characterizing the performance impact of this approach on HPC applications.