Early Termination of Failed HPC Jobs Through Machine and Deep Learning

Early Termination of Failed HPC Jobs Through Machine and Deep Learning
复制标题

通过机器和深度学习提前终止失败的 HPC 作业

DOI:
10.1007/978-3-319-96983-1_12
复制
发表时间:
2018
期刊:
2008 The 28th International Conference on Distributed Computing Systems
影响因子:
--
通讯作者:
T. Ludwig
T. Ludwig
中科院分区:
--
文献类型:
--
作者:
Michal Zasadzinski;V. Muntés;Marc Solé;David Carrera;T. Ludwig

文献摘要

被引文献

相似文献

在超级计算机中,作业失败不仅会造成CPU时间或能量消耗的浪费,而且会降低用户的工作效率。挖掘在数据中心运行期间收集的数据有助于找到解释故障的模式,并可用于预测故障。自动化系统反应,例如,当预测到软件故障时,提前终止作业不仅可以提高可用性并降低操作成本,而且还可以节省管理员和用户的时间。在本文中,我们探讨了一个独特的数据集,包含的拓扑结构,操作指标,和作业调度历史的千万亿次米斯特拉尔超级计算机。我们通过决策树提取最相关的系统特征,决定作业的最终状态。然后,我们成功地训练了一个神经网络预测工作进化的基础上的权力时间序列的节点。最后,我们评估了静态和动态作业终止策略对CPU时间节省的影响。
Failed jobs in a supercomputer cause not only waste in CPU time or energy consumption but also decrease work efficiency of users. Mining data collected during the operation of data centers helps to find patterns explaining failures and can be used to predict them. Automating system reactions, e.g., early termination of jobs, when software failures are predicted does not only increase availability and reduce operating cost, but it also frees administrators’ and users’ time. In this paper, we explore a unique dataset containing the topology, operation metrics, and job scheduler history from the petascale Mistral supercomputer. We extract the most relevant system features deciding on the final state of a job through decision trees. Then, we successfully train a neural net to predict job evolution based on power time series of nodes. Finally, we evaluate the effect on CPU time saving for static and dynamic job termination policies.