Mirage: Towards Low-Interruption Services on Batch GPU Clusters with Reinforcement Learning

Mirage: Towards Low-Interruption Services on Batch GPU Clusters with Reinforcement Learning
复制标题

DOI:
10.1145/3581784.3607042
复制
发表时间:
2023-06
期刊:
SC23: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Qi-Dong Ding;Pengfei Zheng;Shreyas Kudari;S. Venkataraman;Zhao-jie Zhang
Qi-Dong Ding;Pengfei Zheng;Shreyas Kudari;S. Venkataraman;Zhao-jie Zhang
中科院分区:
其他
文献类型:
--
作者:
Qi-Dong Ding;Pengfei Zheng;Shreyas Kudari;S. Venkataraman;Zhao-jie Zhang

文献摘要

相似文献

在使用传统批处理调度器(如Slurm)的GPU集群上,适应长时间运行的深度学习(DL)训练和推理工作是一项挑战。给定固定的时钟时间限制,深度学习研究人员通常需要运行一系列批处理作业,并在过载的机器上经历长时间的中断。这种中断显著降低了在生产环境中部署的服务的研究效率和QoS。为了缓解中断带来的问题,我们提出了一个主动提供程序的设计,并研究了一组统计学习和强化学习(RL)技术,包括随机森林、xgboost、Deep Q-Network和策略梯度。使用来自三个GPU集群的生产作业跟踪,我们使用跟踪的子集训练每个模型,然后使用剩余的验证子集评估它们的通用性。我们介绍Mirage,这是一个与slurm兼容的资源提供程序,它集成了候选ML方法。我们的实验表明,在三个集群的不同负载水平上,Mirage可以减少17-100%的中断,并在零中断的情况下保护23%-76%的作业。
Accommodating long-running deep learning (DL) training and inference jobs is challenging on GPU clusters that use traditional batch schedulers, such as Slurm. Given fixed wall clock time limits, DL researchers usually need to run a sequence of batch jobs and experience long interruptions on overloaded machines. Such interruptions significantly lower the research productivity and QoS for services that are deployed in production. To mitigate the issues from interruption, we propose the design of a proactive provisioner and investigate a set of statistical learning and reinforcement learning (RL) techniques, including random forest, xgboost, Deep Q-Network, and policy gradient. Using production job traces from three GPU clusters, we train each model using a subset of the trace and then evaluate their generality using the remaining validation subset. We introduce Mirage, a Slurm-compatible resource provisioner that integrates the candidate ML methods. Our experiments show that the Mirage can reduce interruption by 17-100% and safe-guard 23%-76% of jobs with zero interruption across varying load levels on the three clusters.