Do Moldable Applications Perform Better on Failure-Prone HPC Platforms?
Do Moldable Applications Perform Better on Failure-Prone HPC Platforms?
复制标题
DOI:
10.1007/978-3-030-10549-5_61
复制
发表时间:
2018-08
期刊:
影响因子:
--
通讯作者:
Valentin Le Fèvre;G. Bosilca;A. Bouteiller;T. Hérault;A. Hori;Y. Robert;J. Dongarra
中科院分区:
文献类型:
--
作者:
Valentin Le Fèvre;G. Bosilca;A. Bouteiller;T. Hérault;A. Hori;Y. Robert;J. Dongarra
This paper compares the performance of different approaches to tolerate failures using checkpoint/restart when executed on large-scale failure-prone platforms. We study (i)applications, which use a constant number of processors throughout execution; (ii)applications, which can use a different number of processors after each restart following a fail-stop error; and (iii)applications, which are moldable applications restricted to use rectangular processor grids (such as many dense linear algebra kernels). For each application type, we compute the optimal number of failures to tolerate before relinquishing the current allocation and waiting until a new resource can be allocated, and we determine the optimal yield that can be achieved. We instantiate our performance model with a realistic applicative scenario and make it publicly available for further usage.