Exploring versioned distributed arrays for resilience in scientific applications

Exploring versioned distributed arrays for resilience in scientific applications
复制标题

DOI:
10.1177/1094342016664796
复制
发表时间:
2017-11
期刊:
The International Journal of High Performance Computing Applications
影响因子:
--
通讯作者:
A. Chien;P. Balaji;N. Dun;A. Fang;H. Fujita;K. Iskra;Z. Rubenstein;Ziming Zheng;J. Hammond;I. Laguna;D. Richards;A. Dubey;B. V. Straalen;M. Hoemmen;M. Heroux;K. Teranishi;A. Siegel
A. Chien;P. Balaji;N. Dun;A. Fang;H. Fujita;K. Iskra;Z. Rubenstein;Ziming Zheng;J. Hammond;I. Laguna;D. Richards;A. Dubey;B. V. Straalen;M. Hoemmen;M. Heroux;K. Teranishi;A. Siegel
中科院分区:
其他
文献类型:
--
作者:
A. Chien;P. Balaji;N. Dun;A. Fang;H. Fujita;K. Iskra;Z. Rubenstein;Ziming Zheng;J. Hammond;I. Laguna;D. Richards;A. Dubey;B. V. Straalen;M. Hoemmen;M. Heroux;K. Teranishi;A. Siegel

文献摘要

被引文献

相似文献

Exascale研究未来HPC系统的项目可靠性挑战。我们提出了全球视图弹性(GVR)系统,便携式弹性库。GVR从全局数组接口的一个子集开始,并添加了新的功能来创建版本,命名版本和计算版本数据。应用程序可以集中在最高效的位置和时间进行版本控制,并独立地为每个应用程序结构进行自定义。该控件具有可移植性,嵌入到应用程序源代码中,表达自然,易于维护。命名多个版本并有效地“部分实现”它们的能力使得基于跨版本或数据结构的“数据切片”的雄心勃勃的向前恢复既容易表达又高效。使用几个大型应用程序(OpenMC,预处理共轭梯度(PCG)求解器,ddcMD和Chombo),我们评估的编程工作,以增加弹性。所需的更改很小(< 2%的代码行),本地化和机器无关,也许最重要的是,不需要软件架构更改。我们还测量了添加GVR版本控制的开销,并表明通常可以实现< 2%的开销。这种开销表明GVR可以在大规模代码中实现,并支持可移植的错误恢复,具有适度的投资和运行时影响。我们的结果来自IBM BG/Q和Cray XC 30实验,证明了可移植性。我们还提出了两个灵活的错误恢复的案例研究,说明如何GVR可以用于多版本回滚恢复,和几个不同的前向恢复计划。GVR的多版本使应用程序能够在具有显著检测延迟的潜在错误(无声数据损坏)中生存下来,并且向前恢复可以使恢复非常有效。我们的研究结果表明,GVR是可扩展的,便携式的,高效的。GVR接口是灵活的,支持各种恢复方案,总的来说,GVR体现了一个缓坡路径,以容忍未来极端规模系统中不断增长的错误率。
Exascale studies project reliability challenges for future HPC systems. We present the Global View Resilience (GVR) system, a library for portable resilience. GVR begins with a subset of the Global Arrays interface, and adds new capabilities to create versions, name versions, and compute on version data. Applications can focus versioning where and when it is most productive, and customize for each application structure independently. This control is portable, and its embedding in application source makes it natural to express and easy to maintain. The ability to name multiple versions and “partially materialize” them efficiently makes ambitious forward-recovery based on “data slices” across versions or data structures both easy to express and efficient. Using several large applications (OpenMC, preconditioned conjugate gradient (PCG) solver, ddcMD, and Chombo), we evaluate the programming effort to add resilience. The required changes are small (< 2% lines of code (LOC)), localized and machine-independent, and perhaps most important, require no software architecture changes. We also measure the overhead of adding GVR versioning and show that overheads < 2% are generally achieved. This overhead suggests that GVR can be implemented in large-scale codes and support portable error recovery with modest investment and runtime impact. Our results are drawn from both IBM BG/Q and Cray XC30 experiments, demonstrating portability. We also present two case studies of flexible error recovery, illustrating how GVR can be used for multi-version rollback recovery, and several different forward-recovery schemes. GVR’s multi-version enables applications to survive latent errors (silent data corruption) with significant detection latency, and forward recovery can make that recovery extremely efficient. Our results suggest that GVR is scalable, portable, and efficient. GVR interfaces are flexible, supporting a variety of recovery schemes, and altogether GVR embodies a gentle-slope path to tolerate growing error rates in future extreme-scale systems.