Efficient Checkpointing with Recompute Scheme for Non-volatile Main Memory

Efficient Checkpointing with Recompute Scheme for Non-volatile Main Memory
复制标题

DOI:
10.1145/3323091
复制
发表时间:
2019-05
期刊:
ACM Transactions on Architecture and Code Optimization (TACO)
影响因子:
--
通讯作者:
Mohammad A. Alshboul;Hussein Elnawawy;Reem Elkhouly;K. Kimura;James Tuck;Yan Solihin
Mohammad A. Alshboul;Hussein Elnawawy;Reem Elkhouly;K. Kimura;James Tuck;Yan Solihin
中科院分区:
其他
文献类型:
--
作者:
Mohammad A. Alshboul;Hussein Elnawawy;Reem Elkhouly;K. Kimura;James Tuck;Yan Solihin

文献摘要

相似文献

未来的主存储器可能包括非易失性存储器。非易失性主内存 (NVMM) 提供了重新思考检查点策略的机会,以便为应用程序提供故障安全性。虽然文献中有许多检查点和日志记录方案,但必须重新审视它们的使用,因为它们会产生较高的执行时间开销以及对 NVMM 的大量额外写入,这可能会显着影响写入耐久性。在本文中,我们提出了一种新颖的基于重新计算的故障安全方法,并演示了其对基于循环的代码的适用性。我们只记录足够的状态来启用重新计算,而不是保持完全一致的日志记录状态。发生故障时,我们的方法通过确定计算的哪些部分未完成并重新计算它们来恢复到一致的状态。实际上,我们的方法消除了保留检查点或日志的需要,从而减少了执行时间开销并提高了 NVMM 写入耐久性,但代价是恢复更复杂。我们在基于 gem5 构建并支持 Intel PMEM 指令扩展的计算机系统模型上,将我们的新方法与五种科学工作负载(包括平铺矩阵乘法)的日志记录和检查点进行了比较。对于平铺矩阵乘法,我们的重新计算方法仅产生 5% 的执行时间开销,而日志记录的开销为 8%,检查点的开销为 207%。此外,重新计算仅增加了 7% 的额外 NVMM 写入,而使用日志记录则增加了 111%,使用检查点则增加了 330%。我们还在真实硬件上进行实验,使我们能够在改变用于计算的线程数量的同时运行工作负载直至完成。这些实验证实了我们基于模拟的观察结果,并提供了重新计算方案和朴素检查点之间的敏感性研究和性能比较。
Future main memory will likely include Non-Volatile Memory. Non-Volatile Main Memory (NVMM) provides an opportunity to rethink checkpointing strategies for providing failure safety to applications. While there are many checkpointing and logging schemes in the literature, their use must be revisited as they incur high execution time overheads as well as a large number of additional writes to NVMM, which may significantly impact write endurance. In this article, we propose a novel recompute-based failure safety approach and demonstrate its applicability to loop-based code. Rather than keeping a fully consistent logging state, we only log enough state to enable recomputation. Upon a failure, our approach recovers to a consistent state by determining which parts of the computation were not completed and recomputing them. Effectively, our approach removes the need to keep checkpoints or logs, thus reducing execution time overheads and improving NVMM write endurance at the expense of more complex recovery. We compare our new approach against logging and checkpointing on five scientific workloads, including tiled matrix multiplication, on a computer system model that was built on gem5 and supports Intel PMEM instruction extensions. For tiled matrix multiplication, our recompute approach incurs an execution time overhead of only 5%, in contrast to 8% overhead with logging and 207% overhead with checkpointing. Furthermore, recompute only adds 7% additional NVMM writes, compared to 111% with logging and 330% with checkpointing. We also conduct experiments on real hardware, allowing us to run our workloads to completion while varying the number of threads used for computation. These experiments substantiate our simulation-based observations and provide a sensitivity study and performance comparison between the Recompute Scheme and Naive Checkpointing.