Scheduling Beyond CPUs for HPC

Scheduling Beyond CPUs for HPC
复制标题

DOI:
10.1145/3307681.3325401
复制
发表时间:
2019-06
期刊:
Proceedings of the 28th International Symposium on High-Performance Parallel and Distributed Computing
影响因子:
--
通讯作者:
Yuping Fan;Z. Lan;Paul M. Rich;W. Allcock;M. Papka;Brian Austin;D. Paul
Yuping Fan;Z. Lan;Paul M. Rich;W. Allcock;M. Papka;Brian Austin;D. Paul
中科院分区:
其他
文献类型:
--
作者:
Yuping Fan;Z. Lan;Paul M. Rich;W. Allcock;M. Papka;Brian Austin;D. Paul

文献摘要

被引文献

相似文献

高性能计算(HPC)正在经历重大变革。新兴的高性能计算应用包括计算密集型和数据密集型应用。为了满足新兴数据密集型应用程序的强烈I/O需求,在生产系统中部署了突发缓冲区。现有的HPC调度器主要以cpu为中心。硬件设备的极端异构性,加上工作负载的变化,迫使调度器在决策时考虑cpu之外的多个资源(例如,突发缓冲区)。在这项研究中,我们提出了一个名为BBSched的多资源调度方案,该方案不仅基于用户的CPU需求,而且基于其他可调度资源(如突发缓冲区)来调度用户的作业。BBSched将调度问题转化为多目标优化(MOO)问题,采用多目标遗传算法快速求解。BBSched生成的多种解决方案使系统管理人员能够探索各种资源之间的潜在权衡,从而更好地利用所有资源。对真实系统工作负载的跟踪驱动模拟表明,与现有方法相比,BBSched将调度性能提高了41%,这表明明确优化cpu以外的多个资源对于HPC调度至关重要。
High performance computing (HPC) is undergoing significant changes. The emerging HPC applications comprise both compute- and data-intensive applications. To meet the intense I/O demand from emerging data-intensive applications, burst buffers are deployed in production systems. Existing HPC schedulers are mainly CPU-centric. The extreme heterogeneity of hardware devices, combined with workload changes, forces the schedulers to consider multiple resources (e.g., burst buffers) beyond CPUs, in decision making. In this study, we present a multi-resource scheduling scheme named BBSched that schedules user jobs based on not only their CPU requirements, but also other schedulable resources such as burst buffer. BBSched formulates the scheduling problem into a multi-objective optimization (MOO) problem and rapidly solves the problem using a multi-objective genetic algorithm. The multiple solutions generated by BBSched enables system managers to explore potential tradeoffs among various resources, and therefore obtains better utilization of all the resources. The trace-driven simulations with real system workloads demonstrate that BBSched improves scheduling performance by up to 41% compared to existing methods, indicating that explicitly optimizing multiple resources beyond CPUs is essential for HPC scheduling.