Learning scheduling algorithms for data processing clusters

Learning scheduling algorithms for data processing clusters
复制标题

DOI:
10.1145/3341302.3342080
复制
发表时间:
2018-10
期刊:
Proceedings of the ACM Special Interest Group on Data Communication
影响因子:
--
通讯作者:
Hongzi Mao;Malte Schwarzkopf;S. Venkatakrishnan;Zili Meng;Mohammad Alizadeh
Hongzi Mao;Malte Schwarzkopf;S. Venkatakrishnan;Zili Meng;Mohammad Alizadeh
中科院分区:
其他
文献类型:
--
作者:
Hongzi Mao;Malte Schwarzkopf;S. Venkatakrishnan;Zili Meng;Mohammad Alizadeh

文献摘要

被引文献

相似文献

在分布式计算集群上有效地调度数据处理作业需要复杂的算法。目前的系统使用简单的,广义的调度和忽略工作负载的特点,因为开发和调整调度策略,为每个工作负载是不可行的。在本文中,我们证明了现代机器学习技术可以自动生成高效的策略。Decima使用强化学习(RL)和神经网络来学习特定于工作负载的调度算法,而无需任何超出高级目标的人工指令,例如最小化平均作业完成时间。然而,现成的RL技术不能处理调度问题的复杂性和规模。为了构建Decima,我们必须为作业的依赖图开发新的表示,设计可扩展的RL模型,并发明RL训练方法来处理连续的随机作业到达。我们在25节点集群上与Spark的原型集成表明,Decima将平均作业完成时间比手动调整的调度算法至少提高了21%,在高集群负载期间实现了高达2倍的改进。
Efficiently scheduling data processing jobs on distributed compute clusters requires complex algorithms. Current systems use simple, generalized heuristics and ignore workload characteristics, since developing and tuning a scheduling policy for each workload is infeasible. In this paper, we show that modern machine learning techniques can generate highly-efficient policies automatically. Decima uses reinforcement learning (RL) and neural networks to learn workload-specific scheduling algorithms without any human instruction beyond a high-level objective, such as minimizing average job completion time. However, off-the-shelf RL techniques cannot handle the complexity and scale of the scheduling problem. To build Decima, we had to develop new representations for jobs' dependency graphs, design scalable RL models, and invent RL training methods for dealing with continuous stochastic job arrivals. Our prototype integration with Spark on a 25-node cluster shows that Decima improves average job completion time by at least 21% over hand-tuned scheduling heuristics, achieving up to 2x improvement during periods of high cluster load.