Salus: Fine-Grained GPU Sharing Primitives for Deep Learning Applications

Salus: Fine-Grained GPU Sharing Primitives for Deep Learning Applications
复制标题

DOI:
--
复制
发表时间:
2019-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Peifeng Yu;Mosharaf Chowdhury
Peifeng Yu;Mosharaf Chowdhury
中科院分区:
其他
文献类型:
--
作者:
Peifeng Yu;Mosharaf Chowdhury

文献摘要

被引文献

相似文献

随着深度学习(DL)应用的激增,GPU计算变得越来越流行。然而,与CPU或网络等传统资源不同,现代GPU本身并不支持细粒度的共享原语。因此,实施诸如分时和抢占之类的公共策略是昂贵的。更糟糕的是,当DL应用不能完全使用GPU的资源时,GPU不能在多个应用之间有效地共享,导致GPU利用率不足。我们提出Salus支持两个GPU共享原语:快速作业切换和内存共享,以实现多个DL应用程序之间的细粒度GPU共享。Salus实现了一个高效的整合执行服务,将GPU暴露给不同的DL应用程序,并通过执行迭代调度和解决相关的内存管理问题来实施细粒度共享。我们表明,这些原语,然后可以用来实现灵活的共享政策,如公平性,优先级和包装的各种用例。Salus与TensorFlow的集成以及对流行DL作业的评估表明,Salus可以将DL训练作业的平均完成时间提高3.19\times $,超参数调整的GPU利用率提高2.38\times $,DL推理应用程序的GPU利用率比不共享GPU提高42\times $,比NVIDIA MPS提高7\times $,并且开销很小。
GPU computing is becoming increasingly more popular with the proliferation of deep learning (DL) applications. However, unlike traditional resources such as CPU or the network, modern GPUs do not natively support fine-grained sharing primitives. Consequently, implementing common policies such as time sharing and preemption are expensive. Worse, when a DL application cannot completely use a GPU's resources, the GPU cannot be efficiently shared between multiple applications, leading to GPU underutilization. We present Salus to enable two GPU sharing primitives: fast job switching and memory sharing, in order to achieve fine-grained GPU sharing among multiple DL applications. Salus implements an efficient, consolidated execution service that exposes the GPU to different DL applications, and enforces fine-grained sharing by performing iteration scheduling and addressing associated memory management issues. We show that these primitives can then be used to implement flexible sharing policies such as fairness, prioritization, and packing for various use cases. Our integration of Salus with TensorFlow and evaluation on popular DL jobs show that Salus can improve the average completion time of DL training jobs by $3.19\times$, GPU utilization for hyper-parameter tuning by $2.38\times$, and GPU utilization of DL inference applications by $42\times$ over not sharing the GPU and $7\times$ over NVIDIA MPS with small overhead.