Characterizing Multi-Instance GPU for Machine Learning Workloads

Characterizing Multi-Instance GPU for Machine Learning Workloads
复制标题

表征机器学习工作负载的多实例 GPU

DOI:
--
复制
发表时间:
2022
期刊:
IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum
影响因子:
--
通讯作者:
Devesh Tiwari
Devesh Tiwari
中科院分区:
--
文献类型:
--
作者:
Baolin Li;V. Gadepally;S. Samsi;Devesh Tiwari

文献摘要

被引文献

相似文献

随着机器学习 (ML) 变得越来越流行,数据中心运营商使用 GPU 等硬件加速器来解决 ML 工作负载的高计算需求。然而,最近的研究表明,用户提交的作业常常未充分利用 GPU 流多处理器 (SM) 内核,从而导致硬件资源浪费。受此启发,GPU 供应商发布了对 GPU 资源共享的软件和硬件支持,例如 A100 Tensor Core GPU 上的 NVIDIA 多实例 GPU (MIG) 技术。在这项工作中,我们使用来自不同应用领域的多种最先进的深度学习 (DL) 模型来表征 A100 GPU MIG 模式操作的性能和能耗。我们的特征揭示了运营支持 MIG 的 GPU 数据中心的宝贵见解。
As machine learning (ML) becomes more and more popular, datacenter operators use hardware accelerators such as GPUs to tackle the high computation demand of ML workloads. However, recent studies show that user-submitted jobs often underutilize the GPU streaming multiprocessor (SM) cores, resulting in hardware resource wastage. Motivated by this observation, GPU vendors have released software and hardware support for GPU resource sharing, for example, the NVIDIA Multi-Instance GPU (MIG) technique on A100 Tensor Core GPUs. In this work, we use several state-of-the-art deep learning (DL) models from various application areas to characterize the performance and energy consumption of the A100 GPU MIG mode operation. Our characterization reveals valuable insights into operating a MIG-enabled GPU datacenter.