Job scheduling for large-scale machine learning clusters

Job scheduling for large-scale machine learning clusters
复制标题

DOI:
10.1145/3386367.3432588
复制
发表时间:
2020-11
期刊:
Proceedings of the 16th International Conference on emerging Networking EXperiments and Technologies
影响因子:
--
通讯作者:
Haoyu Wang;Zetian Liu;Haiying Shen
Haoyu Wang;Zetian Liu;Haiying Shen
中科院分区:
其他
文献类型:
--
作者:
Haoyu Wang;Zetian Liu;Haiying Shen

文献摘要

被引文献

相似文献

随着在现代平台上运行的机器学习(ML)和深度学习(DL)应用程序的快速增长,满足应用程序性能要求至关重要,例如满足最后期限和确保准确性。为此,研究人员为ML集群提出了几种工作分配器。然而,以前提出的并行计算器都没有考虑ML模型并行性,尽管它已经被提出作为一种提高大规模ML和DL作业运行效率的方法。因此,在本文中,我们提出了一个ML作业基于特征的作业调度系统(MLFS)的ML集群运行数据并行和模型并行ML作业。MLFS首先使用一种启发式调度方法,考虑ML作业的空间和时间特征,以确定任务优先级的作业队列排序,以提高作业完成时间(JCT)和精度性能。它使用来自启发式调度方法的数据来训练深度强化学习(RL)模型。在RL模型经过良好的训练后,它将切换到RL方法来自动做出作业调度决策。此外,MLFS具有系统负载控制方法,该方法基于任务优先级从过载服务器选择任务以移动到欠载服务器,并且当系统过载时还智能地移除对期望的准确性性能产生很少或没有改善的任务,以在作业截止日期之前改善JCT和准确性。真实的实验和基于真实的轨迹的大规模仿真表明,与现有的ML作业调度器相比,MLFS将JCT降低了53%,完工时间缩短了52%,准确率提高了64%.我们还开源了我们的代码。
With the rapid proliferation of Machine Learning (ML) and Deep learning (DL) applications running on modern platforms, it is crucial to satisfy application performance requirements such as meeting deadline and ensuring accuracy. To this end, researchers have proposed several job schedulers for ML clusters. However, none of the previously proposed schedulers consider ML model parallelism, though it has been proposed as an approach to increase the efficiency of running large-scale ML and DL jobs. Thus, in this paper, we propose an ML job Feature based job Scheduling system (MLFS) for ML clusters running both data parallelism and model parallelism ML jobs. MLFS first uses a heuristic scheduling method that considers an ML job's spatial and temporal features to determine task priority for job queue ordering in order to improve job completion time (JCT) and accuracy performance. It uses the data from the heuristic scheduling method for training a deep reinforcement learning (RL) model. After the RL model is well trained, it then switches to the RL method to automatically make decisions on job scheduling. Furthermore, MLFS has a system load control method that selects tasks from overloaded servers to move to underloaded servers based on task priority, and also intelligently removes the tasks that generate little or no improvement on the desired accuracy performance when the system is overloaded to improve JCT and accuracy by job deadline. Real experiments and large-scale simulation based on real trace show that MLFS reduces JCT by up to 53% and makespan by up to 52%, and improves accuracy by up to 64% when compared with existing ML job schedulers. We also open sourced our code.