Medea: scheduling of long running applications in shared production clusters

Medea: scheduling of long running applications in shared production clusters
复制标题

DOI:
10.1145/3190508.3190549
复制
发表时间:
2018-04
期刊:
Proceedings of the Thirteenth EuroSys Conference
影响因子:
--
通讯作者:
Panagiotis Garefalakis;Konstantinos Karanasos;P. Pietzuch;Arun Suresh;Sriram Rao
Panagiotis Garefalakis;Konstantinos Karanasos;P. Pietzuch;Arun Suresh;Sriram Rao
中科院分区:
其他
文献类型:
--
作者:
Panagiotis Garefalakis;Konstantinos Karanasos;P. Pietzuch;Arun Suresh;Sriram Rao

文献摘要

被引文献

相似文献

共享生产集群中机器学习、流媒体和延迟敏感型在线应用程序的普及为集群部署者提出了新的挑战。为了优化它们的性能和弹性,这些应用需要通过复杂的约束来精确控制它们的放置,例如,以跨节点组并置或分离其长期运行的容器。在这些应用程序的存在下,集群调度器必须达到全局优化目标,例如最大化部署的应用程序的数量或最小化违反的约束和资源碎片,但不影响短期运行的容器的调度延迟。我们提出了Medea,一个新的集群调度器的长期和短期运行的容器的位置。Medea引入了强大的放置约束和形式语义,以捕获应用程序内和跨应用程序的容器之间的交互。它遵循一种新颖的双调度器设计:(i)对于长时间运行的容器,它应用基于优化的方法,该方法考虑了约束和全局目标;(ii)对于短时间运行的容器,它使用传统的基于任务的调度器来降低放置延迟。在400节点集群上进行评估,我们在Apache Hadoop YARN上实现的Medea实现了长期运行应用程序的放置,与最先进的集群相比,具有显着的性能和弹性优势。
The rise in popularity of machine learning, streaming, and latency-sensitive online applications in shared production clusters has raised new challenges for cluster schedulers. To optimize their performance and resilience, these applications require precise control of their placements, by means of complex constraints, e.g., to collocate or separate their long-running containers across groups of nodes. In the presence of these applications, the cluster scheduler must attain global optimization objectives, such as maximizing the number of deployed applications or minimizing the violated constraints and the resource fragmentation, but without affecting the scheduling latency of short-running containers. We present Medea, a new cluster scheduler designed for the placement of long- and short-running containers. Medea introduces powerful placement constraints with formal semantics to capture interactions among containers within and across applications. It follows a novel two-scheduler design: (i) for long-running containers, it applies an optimization-based approach that accounts for constraints and global objectives; (ii) for short-running containers, it uses a traditional task-based scheduler for low placement latency. Evaluated on a 400-node cluster, our implementation of Medea on Apache Hadoop YARN achieves placement of long-running applications with significant performance and resilience benefits compared to state-of-the-art schedulers.