Extending SLURM for Dynamic Resource-Aware Adaptive Batch Scheduling

Extending SLURM for Dynamic Resource-Aware Adaptive Batch Scheduling
复制标题

扩展 SLURM 以实现动态资源感知自适应批量调度

DOI:
--
复制
发表时间:
2020
期刊:
International Conference on High Performance Computing
影响因子:
--
通讯作者:
M. Gerndt
M. Gerndt
中科院分区:
--
文献类型:
--
作者:
Mohak Chadha;Jophin John;M. Gerndt

文献摘要

被引文献

相似文献

随着电力预算的日益紧张和硬件故障率的增加,未来亿级系统的运行面临着几个挑战。为此,在HPC社区中,通过启用可延展的作业来实现资源意识和适应性已被积极研究。可延展的作业可以在运行时更改其计算资源,并可以显著提高HPC系统的性能。然而,由于流行的并行编程范例(如MPI)的僵硬性质以及对批处理系统中动态资源管理的缺乏支持,可延展性作业在很大程度上没有实现。在本文中,我们扩展了SLURM批处理系统以支持可延展作业的执行和批处理调度。可延展的应用程序是使用称为侵入式MPI的新的自适应并行范例编写的,该范例扩展了MPI标准以支持运行时的资源适应性。我们提出了两种可伸缩的作业调度策略,以支持运行时的性能感知和功耗感知的动态重构决策。我们在SLURM中实现了这些策略,并在生产HPC系统上对它们进行了评估。结果表明,与其他调度策略相比,我们的性能感知调度策略在最长完工时间、平均系统利用率、平均响应时间和等待时间方面都有所改善。此外,我们还使用我们的电力感知策略演示了动态电力走廊管理。
With the growing constraints on power budget and increasing hardware failure rates, the operation of future exascale systems faces several challenges. Towards this, resource awareness and adaptivity by enabling malleable jobs has been actively researched in the HPC community. Malleable jobs can change their computing resources at runtime and can significantly improve HPC system performance. However, due to the rigid nature of popular parallel programming paradigms such as MPI and lack of support for dynamic resource management in batch systems, malleable jobs have been largely unrealized. In this paper, we extend the SLURM batch system to support the execution and batch scheduling of malleable jobs. The malleable applications are written using a new adaptive parallel paradigm called Invasive MPI which extends the MPI standard to support resource-adaptivity at runtime. We propose two malleable job scheduling strategies to support performance-aware and power-aware dynamic reconfiguration decisions at runtime. We implement the strategies in SLURM and evaluate them on a production HPC system. Results for our performance-aware scheduling strategy show improvements in makespan, average system utilization, average response, and waiting times as compared to other scheduling strategies. Moreover, we demonstrate dynamic power corridor management using our power-aware strategy.