Do Moldable Applications Perform Better on Failure-Prone HPC Platforms?

Do Moldable Applications Perform Better on Failure-Prone HPC Platforms?
复制标题

DOI:
10.1007/978-3-030-10549-5_61
复制
发表时间:
2018-08
期刊:
--
影响因子:
--
通讯作者:
Valentin Le Fèvre;G. Bosilca;A. Bouteiller;T. Hérault;A. Hori;Y. Robert;J. Dongarra
Valentin Le Fèvre;G. Bosilca;A. Bouteiller;T. Hérault;A. Hori;Y. Robert;J. Dongarra
中科院分区:
其他
文献类型:
--
作者:
Valentin Le Fèvre;G. Bosilca;A. Bouteiller;T. Hérault;A. Hori;Y. Robert;J. Dongarra

文献摘要

被引文献

相似文献

本文比较了在大规模易发生故障的平台上使用检查点/重启来容忍故障的不同方法的性能。我们研究(i)应用程序,它在整个执行过程中使用恒定数量的处理器;(ii)应用程序,在每次失败停止错误后重新启动后可以使用不同数量的处理器;(iii)应用程序,这是可建模的应用程序,限制使用矩形处理器网格(如许多密集的线性代数核)。对于每种应用程序类型,在放弃当前分配并等待新资源分配之前,我们计算可以容忍的最佳故障数量,并确定可以实现的最佳产量。我们用一个实际的应用场景实例化我们的性能模型,并将其公开供进一步使用。
This paper compares the performance of different approaches to tolerate failures using checkpoint/restart when executed on large-scale failure-prone platforms. We study (i)applications, which use a constant number of processors throughout execution; (ii)applications, which can use a different number of processors after each restart following a fail-stop error; and (iii)applications, which are moldable applications restricted to use rectangular processor grids (such as many dense linear algebra kernels). For each application type, we compute the optimal number of failures to tolerate before relinquishing the current allocation and waiting until a new resource can be allocated, and we determine the optimal yield that can be achieved. We instantiate our performance model with a realistic applicative scenario and make it publicly available for further usage.