Local rollback for resilient MPI applications with application-level checkpointing and message logging

Local rollback for resilient MPI applications with application-level checkpointing and message logging
复制标题

DOI:
10.1016/j.future.2018.09.041
复制
发表时间:
2019-02
期刊:
Future Gener. Comput. Syst.
影响因子:
--
通讯作者:
Nuria Losada;G. Bosilca;A. Bouteiller;P. González;María J. Martín
Nuria Losada;G. Bosilca;A. Bouteiller;P. González;María J. Martín
中科院分区:
其他
文献类型:
--
作者:
Nuria Losada;G. Bosilca;A. Bouteiller;P. González;María J. Martín

文献摘要

被引文献

相似文献

高性能计算(HPC)中通常使用的弹性方法依赖于协调的检查点/重启,即运行应用程序的所有进程的全局回滚。然而,在许多情况下,故障的范围更为局部化,其影响通常仅限于正在使用的资源的子集。因此,全局回滚将导致不必要的开销和能源消耗,因为所有进程,包括那些未受故障影响的进程,都会丢弃其状态并回滚到最后一个检查点,以重复已经完成的计算。用户级故障缓解(ULFM)接口是在消息传递接口(MPI)标准中包含弹性功能的最后一项提议,它支持部署更灵活的恢复策略,包括本地化恢复。通过结合可移植检查点编译器(CPPC)工具ULFM和开放的MPI VProtocol系统级消息日志组件,提出了一种可普遍应用于单程序、多数据(SPMD)应用程序的本地回滚方法。只有失败的进程才能从上一个检查点恢复,而在执行中进一步进展之前的一致性是通过两级消息记录进程实现的。为了进一步优化此方法,点对点通信由Open-MPI VProtocol组件记录,而集体通信则在应用程序级别进行最佳记录-从而将记录协议与特定的集体实施分离。CPPC采用的这种空间协调的协议减少了日志大小,减少了日志存储需求,并总体上减少了对应用程序的弹性影响。
The resilience approach generally used in high-performance computing (HPC) relies on coordinated checkpoint/restart, a global rollback of all the processes that are running the application. However, in many instances, the failure has a more localized scope and its impact is usually restricted to a subset of the resources being used. Thus, a global rollback would result in unnecessary overhead and energy consumption, since all processes, including those unaffected by the failure, discard their state and roll back to the last checkpoint to repeat computations that were already done. The User Level Failure Mitigation (ULFM) interface – the last proposal for the inclusion of resilience features in the Message Passing Interface (MPI) standard – enables the deployment of more flexible recovery strategies, including localized recovery. This work proposes a local rollback approach that can be generally applied to Single Program, Multiple Data (SPMD) applications by combining ULFM, the ComPiler for Portable Checkpointing (CPPC) tool, and the Open MPI VProtocol system-level message logging component. Only failed processes are recovered from the last checkpoint, while consistency before further progress in the execution is achieved through a two-level message logging process. To further optimize this approach point-to-point communications are logged by the Open MPI VProtocol component, while collective communications are optimally logged at the application level—thereby decoupling the logging protocol from the particular collective implementation. This spatially coordinated protocol applied by CPPC reduces the log size, the log memory requirements and overall the resilience impact on the applications.