Modeling the Impact of Checkpoints on Next-Generation Systems

Modeling the Impact of Checkpoints on Next-Generation Systems
复制标题

DOI:
10.1109/msst.2007.24
复制
发表时间:
2007-09
期刊:
24th IEEE Conference on Mass Storage Systems and Technologies (MSST 2007)
影响因子:
--
通讯作者:
R. Oldfield;Sarala Arunagiri;P. Teller;Seetharami R. Seelam;Maria Ruiz Varela;R. Riesen;P. Roth
R. Oldfield;Sarala Arunagiri;P. Teller;Seetharami R. Seelam;Maria Ruiz Varela;R. Riesen;P. Roth
中科院分区:
其他
文献类型:
--
作者:
R. Oldfield;Sarala Arunagiri;P. Teller;Seetharami R. Seelam;Maria Ruiz Varela;R. Riesen;P. Roth

文献摘要

被引文献

相似文献

下一代能力级大规模并行处理(MPP)系统预计将拥有数十万个处理器。对于应用程序驱动的周期性检查点操作,现有技术无法提供可扩展到下一代系统的解决方案。我们证明了这一点,通过使用数学建模来计算这些方法对三个大规模,在生产中,DOE系统和理论petaflop系统上执行的应用程序的性能的影响的下限。我们还调整模型,以研究建议的优化,利用“轻量级”的存储架构和覆盖网络,以克服存储系统的瓶颈。我们的研究结果表明:(1)随着我们接近下一代系统的规模,传统的检查点/重启方法将越来越多地影响应用程序的性能,占总应用程序执行时间的50%以上;(2)虽然我们的替代方法提高了性能,但它有自己的局限性;以及(3)迫切需要新的容错方法,其允许连续计算,而对应用可伸缩性的影响最小。
The next generation of capability-class, massively parallel processing (MPP) systems is expected to have hundreds of thousands of processors. For application-driven, periodic checkpoint operations, the state-of-the-art does not provide a solution that scales to next-generation systems. We demonstrate this by using mathematical modeling to compute a lower bound of the impact of these approaches on the performance of applications executed on three massive-scale, in-production, DOE systems and a theoretical petaflop system. We also adapt the model to investigate a proposed optimization that makes use of "lightweight" storage architectures and overlay networks to overcome the storage system bottleneck. Our results indicate that (1) as we approach the scale of next-generation systems, traditional checkpoint/restart approaches will increasingly impact application performance, accounting for over 50% of total application execution time; (2) although our alternative approach improves performance, it has limitations of its own; and (3) there is a critical need for new approaches to fault tolerance that allow continuous computing with minimal impact on application scalability.