CRUM: Checkpoint-Restart Support for CUDA's Unified Memory

CRUM: Checkpoint-Restart Support for CUDA's Unified Memory
复制标题

DOI:
10.1109/cluster.2018.00047
复制
发表时间:
2018-08
期刊:
2018 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
Rohan Garg;Apoorve Mohan;Michael B. Sullivan;G. Cooperman
Rohan Garg;Apoorve Mohan;Michael B. Sullivan;G. Cooperman
中科院分区:
其他
文献类型:
--
作者:
Rohan Garg;Apoorve Mohan;Michael B. Sullivan;G. Cooperman

文献摘要

被引文献

相似文献

CUDA版本8和Pascal GPU最近引入了统一的虚拟内存(UVM)。较旧的CUDA编程样式类似于用于直接加载和卸载内存段的旧大型内存UNIX应用程序。较新的CUDA程序已经开始利用UVM,其原因与Unix应用程序相同的原因很久以前就转换为假设存在虚拟内存。因此,UVM的检查点已经变得越来越重要,尤其是随着NVIDIA CUDA继续获得更广泛的流行:最新列表中前500个超级计算机中有87个使用NVIDIA GPU,目前每年还具有十个基于NVIDIA的超级计算机。用于跨多个计算机节点的混合CUDA/MPI计算,证明了一种新的可扩展检查点机制CRUM(用于统一存储器的检查点 - 重点)。对UVM的支持对于比GPU上需要更多内存的程序特别有吸引力,因为UVM的替代方案是为了在设备和主机之间直接复制内存。此外,CRUM支持一个快速,分叉的检查点,该检查点主要与CUDA计算与稳定存储中的检查点图像的存储重叠。使用CRUM的运行时开销平均为6%,而分叉检查点的时间被认为是传统同步检查点的40倍。
Unified Virtual Memory (UVM) was recently introduced with CUDA version 8 and the Pascal GPU. The older CUDA programming style is akin to older large-memory UNIX applications which used to directly load and unload memory segments. Newer CUDA programs have started taking advantage of UVM for the same reasons of superior programmability that UNIX applications long ago switched to assuming the presence of virtual memory. Therefore, checkpointing of UVM has become increasing important, especially as NVIDIA CUDA continues to gain wider popularity: 87 of the top 500 supercomputers in the latest listings use NVIDIA GPUs, with a current trend of ten additional NVIDIA-based supercomputers each year. A new scalable checkpointing mechanism, CRUM (Checkpoint-Restart for Unified Memory), is demonstrated for hybrid CUDA/MPI computations across multiple computer nodes. The support for UVM is particularly attractive for programs requiring more memory than resides on the GPU, since the alternative to UVM is for the application to directly copy memory between device and host. Furthermore, CRUM supports a fast, forked checkpointing, which mostly overlaps the CUDA computation with storage of the checkpoint image in stable storage. The runtime overhead of using CRUM is 6% on average, and the time for forked checkpointing is seen to be a factor of up to 40 times less than traditional, synchronous checkpointing.