An Analysis of Resilience Techniques for Exascale Computing Platforms

An Analysis of Resilience Techniques for Exascale Computing Platforms
复制标题

百亿亿次计算平台的弹性技术分析

DOI:
--
复制
发表时间:
2017
期刊:
IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum
影响因子:
--
通讯作者:
H. Siegel
H. Siegel
中科院分区:
--
文献类型:
--
作者:
D. Dauwe;S. Pasricha;A. A. Maciejewski;H. Siegel

文献摘要

被引文献

相似文献

随着大规模高性能计算(HPC)系统中的节点的复杂性和数量的增加,应用经历故障的概率显著增加。随着在HPC系统上执行的应用程序的计算需求的增加,预测表明,在兆级大小的系统上执行的应用程序可能会以低至几分钟的平均故障间隔时间(MTBF)运行。近年来,已经提出了一些在极端规模的系统中实现故障恢复的策略。然而,很少有研究提供这些弹性技术的性能比较。这项工作提供了一个比较四个国家的最先进的HPC弹性技术,正在考虑使用在exascale系统。我们探索的行为,每个弹性技术下模拟执行的一组不同的应用程序在通信行为和内存使用。我们研究了每种弹性技术在应用程序规模从今天被认为是大型的应用程序扩展到千兆级应用程序时的行为。我们进一步研究了大规模系统从与每个弹性技术相关的开销以及在故障发生时继续执行所需的应用程序计算中经历的性能下降。使用这些分析的结果,我们研究如何在exascale系统上的应用程序的性能可以通过允许系统选择最佳的弹性技术用于在特定于应用程序的方式,根据每个应用程序的执行特性来提高。
With the increase in the complexity and number of nodes in large-scale high performance computing (HPC) systems, the probability of applications experiencing failures has increased significantly. As the computational demands of applications that execute on HPC systems increase, projections indicate that applications executing on exascale-sized systems are likely to operate with a mean time between failures (MTBF) of as little as a few minutes. A number of strategies for enabling fault resilience in systems of extreme sizes have been proposed in recent years. However, few studies provide performance comparisons for these resilience techniques. This work provides a comparison of four state-of-the-art HPC resilience techniques that are being considered for use in exascale systems. We explore the behavior of each resilience technique under simulated execution of a diverse set of applications varying in communication behavior and memory use. We examine how each resilience technique behaves as application size scales from what is considered large today through to exascale-sized applications. We further study the performance degradation that a large-scale system experiences from the overhead associated with each resilience technique as well as the application computation needed to continue execution when a failure occurs. Using the results from these analyses, we examine how application performance on exascale systems can be improved by allowing the system to select the optimal resilience technique for use in an application-specific manner, depending upon each application's execution characteristics.