Design and Implementation for Checkpointing of Distributed Resources Using Process-Level Virtualization

Design and Implementation for Checkpointing of Distributed Resources Using Process-Level Virtualization
复制标题

使用进程级虚拟化的分布式资源检查点的设计和实现

DOI:
10.1109/cluster.2016.55
复制
发表时间:
2016
期刊:
2016 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
G. Cooperman
G. Cooperman
中科院分区:
--
文献类型:
--
作者:
K. Arya;Rohan Garg;A. Y. Polyakov;G. Cooperman

文献摘要

被引文献

相似文献

系统级检查点重启是高性能计算中长时间运行作业的关键技术。然而,只有两种检查点MPI应用程序的方法在今天继续广泛使用。一种方法是将基于内核模块的BLCR与特定于所使用的MPI实现的MPI检查点重启服务结合使用。不幸的是,这缺乏对一些重要的Linux系统服务的支持,如SysV IPC(例如,共享内存对象)。第二种方法是使用原始的2009 DMTCP实现(本文称为DMTCP-09)进行透明的系统级检查点设置。不幸的是,DMTCP-09缺乏对MPI在现代批处理环境中发现的许多必要功能的检查点支持。其中包括:ssh、InfiniBand网络、进程迁移(在不同的集群节点上重新启动MPI应用程序)以及重新启动时修改的文件路径前缀(通常是由于更改当前目录、挂载点、库路径等)。这项工作提出了DMTCP-PV,一个新的用户空间透明的检查点系统的基础上的概念,进程虚拟化。这种方法分别对每个本地或分布式子系统的状态进行建模,同时将其与核心检查点引擎解耦。通过分离这些关注点,领域专家可以将检查点扩展到一个新的领域,而无需了解核心检查点引擎。这使得DMTCP-PV能够解决上述缺陷和许多其他缺陷。结果表明,DMTCP-PV的运行时开销一般小于1%,检查点时间主要是由一个图像文件写入稳定存储的时间。
System-level checkpoint-restart is a critical technology for long-running jobs in high-performance computing. Yet, only two approaches to checkpointing MPI applications continue to survive in wide use today. One approach is to use the kernel module-based BLCR in combination with an MPI checkpoint-restart service particular to the MPI implementation in use. Unfortunately, this lacks support for some important Linux system services such as SysV IPC (e.g., shared memory objects). A second approach has been to use the original 2009 DMTCP implementation (herein referred to as DMTCP-09) for transparent, system-level checkpointing. Unfortunately, DMTCP-09 lacked support for checkpointing many of the necessary features found by MPI in a modern batch environment. These include: ssh, the InfiniBand network, process migration (restarting an MPI application on different cluster nodes), and modified file path prefixes on restart (typically due to a changing current directory, mount points, library paths, etc.). This work presents DMTCP-PV, a new user-space transparent checkpointing system based on the concept of process virtualization. This approach separately models the state of each local or distributed subsystem while decoupling it from the core checkpointing engine. By separating these concerns, a domain expert can extend checkpointing into a new domain without any knowledge of the core checkpointing engine. This allowed DMTCP-PV to address the deficiencies noted above and many others. It is shown that the runtime overhead of DMTCP-PV is generally less than 1%, and the checkpointing time is dominated by the time to write an image file to stable storage.