Partitioned Parallel Job Scheduling for Extreme Scale Computing

Partitioned Parallel Job Scheduling for Extreme Scale Computing
复制标题

超大规模计算的分区并行作业调度

DOI:
10.1007/978-3-642-35867-8_9
复制
发表时间:
2012
期刊:
2014 22nd Euromicro International Conference on Parallel, Distributed, and Network-Based Processing
影响因子:
--
通讯作者:
Seetharami R. Seelam
Seetharami R. Seelam
中科院分区:
--
文献类型:
--
作者:
David Brelsford;G. Chochia;Nathan Falk;Kailash Marthi;Ravindra Sure;N. Bobroff;L. Fong;Seetharami R. Seelam

文献摘要

被引文献

相似文献

最近在构建极限计算系统方面取得的成功给作业调度设计带来了新的挑战,以支持可以执行数百万个并发任务的集群大小。我们表明,对于这些极端规模的集群,集中式调度程序的资源需求可能会超出容量或限制调度程序良好执行的能力。本文介绍了分区调度,这是一种集中式和分布式混合的方法,其中计算节点集中分配给作业,而任务到本地节点资源的分配随后在分配的作业节点上执行。这减少了中央调度程序的内存和处理增长,并通过允许在作业节点并行完成操作来改进调度时间的扩展行为。当本地资源分配必须分配给所有其他作业节点时,分区方法会牺牲中央处理来增加网络通信。因此,我们引入了改善通信的功能,例如利用高速集群网络的管道。新系统针对具有 496 个节点、每个节点 128 个任务的集群上最多 50K 任务的作业进行了评估。分区调度方法被证明可以减少中央处理器的处理器和内存使用,并将作业调度和作业分派时间提高一个数量级。
Recent success in building extreme computing systems poses new challenges in job scheduling design to support cluster sizes that can execute million’s of concurrent tasks. We show that for these extreme scale clusters the resource demand at a centralized scheduler can exceed the capacity or limit the ability of the scheduler to perform well. This paper introduces partitioned scheduling, a hybrid centralized and distributed approach in which compute nodes are assigned to the job centrally, while task to local node resources assignments are performed subsequently at the assigned job nodes. This reduces the memory and processing growth at the central scheduler, and improves the scaling behavior of scheduling time by enabling operations to be done in parallel at the job nodes. When local resource assignments must be distributed to all other job nodes, the partitioned approach trades central processing for increased network communications. Thus, we introduce features that improve communications such as pipelining that leverage the presence of the high speed cluster network. The new system is evaluated for jobs with up to 50K tasks on clusters with 496 nodes and 128 tasks per node. The partitioned scheduling approach is demonstrated to reduce processor and memory usage at the central processor and improve job scheduling and job dispatching times up to an order of magnitude.