Fault-tolerant communication runtime support for data-centric programming models

Fault-tolerant communication runtime support for data-centric programming models
复制标题

针对以数据为中心的编程模型的容错通信运行时支持

DOI:
10.1109/hipc.2010.5713195
复制
发表时间:
2010
期刊:
2010 International Conference on High Performance Computing
影响因子:
--
通讯作者:
S. Song
S. Song
中科院分区:
--
文献类型:
--
作者:
Abhinav Vishnu;H. V. Dam;W. D. Jong;P. Balaji;S. Song

文献摘要

被引文献

相似文献

当今世界上最大的超级计算机由数十万个处理核心和更多其他硬件组件组成。在这样的规模下,硬件故障是司空见惯的,需要具有容错能力的软件系统。虽然有不同的容错模型可用,但大多数模型都侧重于允许计算过程在故障中存活下来。另一方面,我们最近开始研究以数据为中心的编程模型(如分区全局地址空间(PGAS)模型)的容错技术。以数据为中心的模型的主要区别是计算和数据局部性的分离。也就是说,数据放置与正在执行的进程分离,允许我们将进程故障(托管进程的物理节点失效)与数据故障(托管数据的物理节点失效)分开查看。在本文中,我们通过使用全局阵列及其通信系统ARMCI设计和实现了一个容错的单边通信运行时框架,从而向以数据为中心的容错迈出了第一步。该框架包括容错进程管理器;低开销和网络辅助的远程节点故障检测模块;非数据移动的集体通信原语;故障语义和单边通信运行时系统的错误或代码。我们的性能评估表明,与最先进的设计相比,该框架的开销很小,并为PGAS模型提供了一个基本的容错框架。
The largest supercomputers in the world today consist of hundreds of thousands of processing cores and many more other hardware components. At such scales, hardware faults are a commonplace, necessitating fault-resilient software systems. While different fault-resilient models are available, most focus on allowing the computational processes to survive faults. On the other hand, we have recently started investigating fault resilience techniques for data-centric programming models such as the partitioned global address space (PGAS) models. The primary difference in data-centric models is the decoupling of computation and data locality. That is, data placement is decoupled from the executing processes, allowing us to view process failure (a physical node hosting a process is dead) separately from data failure (a physical node hosting data is dead). In this paper, we take a first step toward data-centric fault resilience by designing and implementing a fault-resilient, onesided communication runtime framework using Global Arrays and its communication system, ARMCI. The framework consists of a fault-resilient process manager; low-overhead and networkassisted remote-node fault detection module; non-data-moving collective communication primitives; and failure semantics and err or codes for one-sided communication runtime systems. Our performance evaluation indicates that the framework incurs little ov erhead compared to state-of-the-art designs and provides a fundamental framework of fault resiliency for PGAS models.