Reverse computation for rollback-based fault tolerance in large parallel systems

Reverse computation for rollback-based fault tolerance in large parallel systems
复制标题

DOI:
10.1007/s10586-013-0277-4
复制
发表时间:
2014-06-01
影响因子:
4.4
通讯作者:
Park, Alfred J.
Park, Alfred J.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Perumalla, Kalyan S.;Park, Alfred J.

文献摘要

被引文献

相似文献

反向计算在此被作为一个重要的未来方向提出,用于应对在大规模并行计算集群平台上进行容错执行的挑战。随着并行作业规模的增大,传统的检查点方法面临可扩展性问题,从计算速度减慢到检查点在持久存储中的高度拥塞不等。反向计算能够克服这些问题,并且也更适合在具有更小、更廉价或节能的存储器和文件系统的新型架构上进行并行计算。通过一个粒子(理想气体)模拟扩展到65536个处理器核心和950个加速器(GPU)的详细性能数据,给出了在大型系统中反向计算可行性的初步证据。当节点依靠其主机处理器/内存来容忍其加速器的故障时,观察到反向计算相对于检查点方案能带来非常大的收益。通过对缓存未命中率、TLB未命中率和内存使用等指标的测量,对反向计算和检查点进行比较,结果表明,在新兴架构中,作为一种未来可供选择的方法,反向计算不容忽视。
Reverse computation is presented here as an important future direction in addressing the challenge of fault tolerant execution on very large cluster platforms for parallel computing. As the scale of parallel jobs increases, traditional checkpointing approaches suffer scalability problems ranging from computational slowdowns to high congestion at the persistent stores for checkpoints. Reverse computation can overcome such problems and is also better suited for parallel computing on newer architectures with smaller, cheaper or energy-efficient memories and file systems. Initial evidence for the feasibility of reverse computation in large systems is presented with detailed performance data from a particle (ideal gas) simulation scaling to 65,536 processor cores and 950 accelerators (GPUs). Reverse computation is observed to deliver very large gains relative to checkpointing schemes when nodes rely on their host processors/memory to tolerate faults at their accelerators. A comparison between reverse computation and checkpointing with measurements such as cache miss ratios, TLB misses and memory usage indicates that reverse computation is hard to ignore as a future alternative to be pursued in emerging architectures.