Interconnect agnostic checkpoint/restart in open MPI

Interconnect agnostic checkpoint/restart in open MPI
复制标题

开放 MPI 中的互连不可知检查点/重启

DOI:
10.1145/1551609.1551619
复制
发表时间:
2009
期刊:
2007 IEEE International Parallel and Distributed Processing Symposium
影响因子:
--
通讯作者:
A. Lumsdaine
A. Lumsdaine
中科院分区:
--
文献类型:
--
作者:
Joshua Hursey;T. Mattox;A. Lumsdaine

文献摘要

被引文献

相似文献

如果要利用当前和未来的高性能计算(HPC)系统,长期运行的大规模HPC应用必须能够容忍不可避免的故障。消息传递接口(MPI)级别的透明检查点/重启容错对于不希望重构其代码的HPC应用程序开发人员来说是一个有吸引力的选择。从历史上看,提供此选项的MPI实现一直在努力提供全方位的互连支持,特别是共享内存支持。本文提出了一种新的方法来实现检查点/重启协调算法,允许MPI实现的检查点/重启互连不可知。该方法允许在一组互连上对应用设置检查点(例如,InfiniBand和共享存储器)并且可以用不同的互连集合重新启动(例如,Myrinet和共享内存或以太网)。通过将网络互连细节与检查点/重启协调算法分离,我们允许HPC应用程序响应群集环境中的变化,例如由于交换机故障导致的互连不可用、在现有机器上重新负载平衡或迁移到具有不同互连集的不同机器。我们目前的结果表征这种方法对HPC应用程序的性能影响。
Long running High Performance Computing (HPC) applications at scale must be able to tolerate inevitable faults if they are to harness current and future HPC systems. Message Passing Interface (MPI) level transparent checkpoint/restart fault tolerance is an appealing option to HPC application developers that do not wish to restructure their code. Historically, MPI implementations that provided this option have struggled to provide a full range of interconnect support, especially shared memory support. This paper presents a new approach for implementing checkpoint/restart coordination algorithms that allows the MPI implementation of checkpoint/restart to be interconnect agnostic. This approach allows an application to be checkpointed on one set of interconnects (e.g., InfiniBand and shared memory) and be restarted with a different set of interconnects (e.g., Myrinet and shared memory or Ethernet). By separating the network interconnect details from the checkpoint/restart coordination algorithm we allow the HPC application to respond to changes in the cluster environment such as interconnect unavailability due to switch failure, re-load balance on an existing machine, or migrate to a different machine with a different set of interconnects. We present results characterizing the performance impact of this approach on HPC applications.