MPI Stages: Checkpointing MPI State for Bulk Synchronous Applications

MPI Stages: Checkpointing MPI State for Bulk Synchronous Applications
复制标题

MPI 阶段:批量同步应用程序的 MPI 状态检查点

DOI:
10.1145/3236367.3236385
复制
发表时间:
2018
期刊:
Proceedings of the 25th European MPI Users' Group Meeting
影响因子:
--
通讯作者:
M. Emani
M. Emani
中科院分区:
--
文献类型:
--
作者:
Nawrin Sultana;A. Skjellum;I. Laguna;M. Farmer;K. Mohror;M. Emani

文献摘要

被引文献

相似文献

当MPI程序遇到故障时,最常见的恢复方法是从以前的检查点重新启动所有过程,并重新征服整个工作。流程必须从程序的开头开始,以及新的替换过程 - 这会为实时流程提供不必要的开销。 “ MPI阶段”的概念。在与应用程序状态下,在单独的检查点中保存了内部MPI状态。实时过程仅在主循环中回滚几个迭代,而不是回到程序的开头,而替换失败的过程重新启动并重新融合,从而在本文中,实现更快的失败恢复与大规模同步应用程序和检查点/重新启动。序列化和应对MPI对象的序列化。这包括MPI阶段。
When an MPI program experiences a failure, the most common recovery approach is to restart all processes from a previous checkpoint and to re-queue the entire job. A disadvantage of this method is that, although the failure occurred within the main application loop, live processes must start again from the beginning of the program, along with new replacement processes---this incurs unnecessary overhead for live processes. To avoid such overheads and concomitant delays, we introduce the concept of "MPI Stages." MPI Stages saves internal MPI state in a separate checkpoint in conjunction with application state. Upon failure, both MPI and application state are recovered, respectively, from their last synchronous checkpoints and continue without restarting the overall MPI job. Live processes roll back only a few iterations within the main loop instead of rolling back to the beginning of the program, while a replacement of failed process restarts and reintegrates, thereby achieving faster failure recovery. This approach integrates well with large-scale, bulk synchronous applications and checkpoint/restart. In this paper, we identify requirements for production MPI implementations to support state checkpointing with MPI Stages, which includes capturing and managing internal MPI state and serializing and deserializing user handles to MPI objects. We evaluate our fault tolerance approach with a proof-of-concept prototype MPI implementation that includes MPI Stages. We demonstrate its functionality and performance using LULESH and microbenchmarks. Our results show that MPI Stages reduces the recovery time by 13× for LULESH in comparison to checkpoint/restart.