Comparing different approaches for Incremental Checkpointing : The Showdown

Comparing different approaches for Incremental Checkpointing : The Showdown
复制标题

比较增量检查点的不同方法:摊牌

DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
Paul H. Hargrove
Paul H. Hargrove
中科院分区:
--
文献类型:
--
作者:
M. Vasavada;F. Mueller;Paul H. Hargrove

文献摘要

被引文献

相似文献

高性能计算(HPC)中核心和节点数量的快速增加使得千万亿次计算成为现实,亿亿次计算即将到来。利用这种计算能力提出了一个挑战,因为系统可靠性随着给定单个单元可靠性的建筑部件的增加而恶化。今天的高端HPC安装要求应用程序执行检查点,如果他们想大规模运行,以便在运行超过几小时或几天的故障可以通过从最后一个检查点重新启动来处理。然而,这样的检查点导致高开销,这是由于所有节点经常同时写入并行文件系统(PFS),这降低了这样的系统在吞吐量计算方面的生产力。最近关于检查点/重启(C/R)的工作表明,增量C/R技术可以减少在检查点写入的数据量,从而减少总体C/R开销和对PFS的影响。这项工作的贡献是双重的。首先,它提出了两个内存管理方案,使增量检查点的设计和实现。我们描述了独特的方法,增量检查点,不需要内核修补在一种情况下,只需要最小的内核扩展在其他情况下。这项工作是在最新的伯克利实验室检查点重启(BLCR)中进行的,作为即将发布的一部分。其次,我们评估了这两种方案的系统开销为单节点微基准和多节点集群工作负载。简而言之,这项工作是页面写位(WB)保护和脏位(DB)页面跟踪作为支持增量检查点的硬件手段之间的最后摊牌。我们的研究结果表明,在几乎所有的测试中,DB方法比WB方法节省。此外,DB具有显著减少内核活动的潜力,这对于主动容错具有最大的相关性,其中如果基于DB的实时迁移将进程从即将发生故障的硬件中移走,则可以规避即时故障。
The rapid increase in the number of cores and nodes in high performance computing (HPC) has made petascale computing a reality with exascale on the horizon. Harnessing such computational power presents a challenge as system reliability deteriorates with the increase of building components of a given single-unit reliability. Today’s high-end HPC installations require applications to perform checkpointing if they want to run at scale so that failures during runs over hours or days can be dealt with by restarting from the last checkpoint. Yet, such checkpointing results in high overheads due to often simultaneous writes of all nodes to the parallel file system (PFS), which reduces the productivity of such systems in terms of throughput computing. Recent work on checkpoint/restart (C/R) has shown that incremental C/R techniques can reduce the amount of data written at checkpoints and thus the overall C/R overhead and impact on the PFS. The contributions of this work are twofold. First, it presents the design and implementation of two memory management schemes that enable incremental checkpointing. We describe unique approaches to incremental checkpointing that do not require kernel patching in one case and only require minimal kernel extensions in the other case. The work is carried out within the latest Berkeley Labs Checkpoint Restart (BLCR) as part of an upcoming release. Second, we evaluate the two schemes in terms of their system overhead for single-node microbenchmarks and multi-node cluster workloads. In short, this work is the final showdown between page write bit (WB) protection and dirty bit (DB) page tracking as a hardware means to support incremental checkpointing. Our results show savings of the DB approach over WB approach in almost all the tests. Further, DB has the potential of a significant reduction in kernel activity, which is of utmost relevance for proactive fault tolerance where an immanent fault can be circumvented if DB-based live migrations moves a process away from hardware about to fail.