Fault tolerance for remote memory access programming models

Fault tolerance for remote memory access programming models
复制标题

远程内存访问编程模型的容错

DOI:
--
复制
发表时间:
2014
期刊:
IEEE International Symposium on High-Performance Parallel Distributed Computing
影响因子:
--
通讯作者:
T. Hoefler
T. Hoefler
中科院分区:
--
文献类型:
--
作者:
Maciej Besta;T. Hoefler

文献摘要

被引文献

相似文献

远程内存访问(RMA)是一种新兴的高性能计算机和嵌入式系统编程机制。然而,关于基于风险管理的应用程序和系统的复原力计划的工作很少。在本文中,我们分析了RMA的容错性,并表明它是从根本上不同于针对消息传递(MP)模型的弹性机制。我们设计了一个模型的推理RMA的容错,解决平面和层次化的硬件。我们使用这个模型来构建几个高度可扩展的机制,提供高效的低开销的内存中的检查点,透明的远程内存访问日志记录,并透明恢复失败的进程的计划。我们的协议考虑到每个核心的内存减少,这是未来亿兆级机器的主要特征之一。我们的容错方案的实施需要微不足道的额外开销。我们的可靠性模型表明,内存中的检查点和日志记录提供了高弹性。该研究为RMA提供了高度可扩展的弹性机制,填补了容错和新兴RMA编程模型之间的研究空白。
Remote Memory Access (RMA) is an emerging mechanism for programming high-performance computers and datacenters. However, little work exists on resilience schemes for RMA-based applications and systems. In this paper we analyze fault tolerance for RMA and show that it is fundamentally different from resilience mechanisms targeting the message passing (MP) model. We design a model for reasoning about fault tolerance for RMA, addressing both flat and hierarchical hardware. We use this model to construct several highly-scalable mechanisms that provide efficient low-overhead in-memory checkpointing, transparent logging of remote memory accesses, and a scheme for transparent recovery of failed processes. Our protocols take into account diminishing amounts of memory per core, one of the major features of future exascale machines. The implementation of our fault-tolerance scheme entails negligible additional overheads. Our reliability model shows that in-memory checkpointing and logging provide high resilience. This study enables highly-scalable resilience mechanisms for RMA and fills a research gap between fault tolerance and emerging RMA programming models.