Euro-Par 2019: Parallel Processing Workshops - Euro-Par 2019 International Workshops, Göttingen, Germany, August 26-30, 2019, Revised Selected Papers

Euro-Par 2019: Parallel Processing Workshops - Euro-Par 2019 International Workshops, Göttingen, Germany, August 26-30, 2019, Revised Selected Papers
复制标题

Euro-Par 2019:并行处理研讨会 - Euro-Par 2019 国际研讨会,德国哥廷根,2019 年 8 月 26-30 日,修订后的精选论文

DOI:
10.1007/978-3-030-48340-1_53
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Baird M
Baird M
中科院分区:
--
文献类型:
--
作者:
Baird M

文献摘要

相似文献

本文提出了一种新的方法来检查点MPI应用程序,使用长时间运行的CUDA内核。可以对驻留在GPU上的数据进行快照,而无需等待内核完成。所提出的技术是在最先进的高性能容错库FTI的背景下实现的。因此,我们得到了一个优雅的解决方案,开发弹性MPI应用程序,其中GPU内核运行时间长于硬件故障之间的平均时间。我们详细描述了我们如何检查点/重新启动协作MPI-CUDA应用程序,我们提供了一个初步评估所提出的方法使用利弗莫尔非结构拉格朗日显式冲击流体动力学(LULESH)应用程序作为案例研究。
This paper proposes a new approach to checkpointing MPI applications that use long-running CUDA kernels. It becomes possible to take snapshots of data residing on the GPUs without waiting for kernels to complete. The proposed technique is implemented in the context of the state of the art high performance fault tolerance library FTI. As a result we get an elegant solution to the problem of developing resilient MPI applications where GPU kernels run longer than the mean time between hardware failures. We describe in detail how we checkpoint/restart collaborative MPI-CUDA applications, and we provide an initial evaluation of the proposed approach using the Livermore Unstructured Lagrangian Explicit Shock Hydrodynamics (LULESH) application as a case study.