MARBLE: A Multi-GPU Aware Job Scheduler for Deep Learning on HPC Systems

MARBLE: A Multi-GPU Aware Job Scheduler for Deep Learning on HPC Systems
复制标题

DOI:
10.1109/ccgrid49817.2020.00-66
复制
发表时间:
2020-05
期刊:
2020 20th IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGRID)
影响因子:
--
通讯作者:
Jingoo Han;M. M. Rafique-M.;Luna Xu;A. Butt;Seung-Hwan Lim;Sudharshan S. Vazhkudai
Jingoo Han;M. M. Rafique-M.;Luna Xu;A. Butt;Seung-Hwan Lim;Sudharshan S. Vazhkudai
中科院分区:
其他
文献类型:
--
作者:
Jingoo Han;M. M. Rafique-M.;Luna Xu;A. Butt;Seung-Hwan Lim;Sudharshan S. Vazhkudai

文献摘要

被引文献

相似文献

深度学习已成为解决复杂科学问题的重要工具。然而,管理与DL相关的多维大规模数据,特别是在现代超级计算机中现有的多个图形处理单元(GPU)之上,带来了巨大的挑战。此外,与现有研究相比,最新的高性能计算(HPC)架构在训练吞吐量方面带来了不同的性能趋势。由于CPU到GPU的快速连接,现有的下行优化(如更大的批处理大小和GPU位置感知调度)对提高下行训练吞吐量性能几乎没有影响。此外,在多个GPU上进行的DL培训可进行子线性扩展。因此,简单地向系统添加更多的GPU是无效的。为此,我们设计了一个首创的作业调度器MARBLE,该调度器在节点级考虑了GPU的非线性可伸缩性,为作业调度每个节点的合适数量的GPU。通过共享具有多个DL作业的节点上的GPU资源,Marble避免了当前HPC系统上的多GPU DL培训中的低GPU利用率。我们在Summit超级计算机上的综合评估表明,与流行的平台负载共享设施(LSF)调度器相比,Marble能够将DL训练性能提高高达48.3%。与最先进的DL调度程序Optimus相比,大理石将作业完成时间减少了高达47%。
Deep learning (DL) has become a key tool for solving complex scientific problems. However, managing the multi-dimensional large-scale data associated with DL, especially atop extant multiple graphics processing units (GPUs) in modern supercomputers poses significant challenges. Moreover, the latest high-performance computing (HPC) architectures bring different performance trends in training throughput compared to the existing studies. Existing DL optimizations such as larger batch size and GPU locality-aware scheduling have little effect on improving DL training throughput performance due to fast CPU-to-GPU connections. Additionally, DL training on multiple GPUs scales sublinearly. Thus, simply adding more GPUs to a system is ineffective. To this end, we design MARBLE, a first-of-its-kind job scheduler, which considers the non-linear scalability of GPUs at the intra-node level to schedule an appropriate number of GPUs per node for a job. By sharing the GPU resources on a node with multiple DL jobs, MARBLE avoids low GPU utilization in current multi-GPU DL training on HPC systems. Our comprehensive evaluation in the Summit supercomputer shows that MARBLE is able to improve DL training performance by up to 48.3% compared to the popular Platform Load Sharing Facility (LSF) scheduler. Compared to the state-of-the-art of DL scheduler, Optimus, MARBLE reduces the job completion time by up to 47%.