Characterizing Multi-Instance GPU for Machine Learning Workloads
Characterizing Multi-Instance GPU for Machine Learning Workloads
复制标题
表征机器学习工作负载的多实例 GPU
DOI:
--
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Devesh Tiwari
中科院分区:
文献类型:
--
作者:
Baolin Li;V. Gadepally;S. Samsi;Devesh Tiwari
As machine learning (ML) becomes more and more popular, datacenter operators use hardware accelerators such as GPUs to tackle the high computation demand of ML workloads. However, recent studies show that user-submitted jobs often underutilize the GPU streaming multiprocessor (SM) cores, resulting in hardware resource wastage. Motivated by this observation, GPU vendors have released software and hardware support for GPU resource sharing, for example, the NVIDIA Multi-Instance GPU (MIG) technique on A100 Tensor Core GPUs. In this work, we use several state-of-the-art deep learning (DL) models from various application areas to characterize the performance and energy consumption of the A100 GPU MIG mode operation. Our characterization reveals valuable insights into operating a MIG-enabled GPU datacenter.