Diagnosing the Interference on CPU-GPU Synchronization Caused by CPU Sharing in Multi-Tenant GPU Clouds

Diagnosing the Interference on CPU-GPU Synchronization Caused by CPU Sharing in Multi-Tenant GPU Clouds
复制标题

DOI:
10.1109/ipccc51483.2021.9679439
复制
发表时间:
2021-10
期刊:
2021 IEEE International Performance, Computing, and Communications Conference (IPCCC)
影响因子:
--
通讯作者:
Youssef Elmougy;Weiwei Jia;Xiaoning Ding;Jianchen Shan
Youssef Elmougy;Weiwei Jia;Xiaoning Ding;Jianchen Shan
中科院分区:
其他
文献类型:
--
作者:
Youssef Elmougy;Weiwei Jia;Xiaoning Ding;Jianchen Shan

文献摘要

被引文献

相似文献

通过成熟的GPU虚拟化技术启用的GPU加速云已成为高性能计算和机器学习工作负载的最吸引人的平台。但是,众所周知,建立多租户GPU云是可以共享资源(例如CPU和GPU)的挑战。一个知名且经过深思熟虑的原因是,当分享GPU时,工作负载的性能隔离不良和GPU利用率较低。但是,很少关注另一个基本问题,但在研究的问题下:在GPU实例之间共享CPU如何影响工作负载的性能?针对此问题,本文进行了实验,以衡量CPU共享干扰下的性能放缓和VGPU利用率下降。结果表明,由于CPU共享,GPU的工作负载遭受了较差和不可预测的性能和大量的VGPU的不足。我们发现,这种干扰是CPU-GPU相互作用的特征与共享VCPU的特殊行为之间的复杂相互作用的结果:VCPU不连续性。为了诊断VCPU不连续性如何引起干扰,纸张利用Nvidia nsight Systems进行细粒度分析,并具有以下发现:1)VCPU不连续性导致效率低下的CPU-GPU同步; 2)VCPU不连续性将任务延迟到VGPU; 3)基于轮询的CPU-GPU同步比基于阻止基于CPU-GPU同步的干扰更大。 4)具有频繁任务卸载和同步的GPU工作负载更加脆弱。根据研究结果,本文提出了一种新颖的轮询 - 然后阻塞CPU-GPU同步原始。评估表明,它可以提高4.2倍的性能。
The GPU-accelerated cloud, enabled by maturing GPU virtualization techniques, has become the most attractive platform for high-performance computing and machine learning workloads. However, it is notoriously challenging to build the multi-tenant GPU cloud where resources, like CPUs and GPUs, can be shared. One well-known and heavily studied reason is that workloads suffer from poor performance isolation and low GPU utilization when GPUs are shared. But little attention has been paid to another fundamental yet under studied problem: how sharing CPUs among GPU instances could affect the workload performance?Targeting this problem, the paper conducts experiments to measure the performance slowdown and vGPU utilization decrease under interference from CPU sharing. The results show that GPU workloads suffer from poor and unpredictable performance and heavy vGPU under-utilization because of CPU sharing. We find that such interference is the result of the complex interplay between the characteristics of CPU-GPU interactions and the special behavior of shared vCPUs: vCPU discontinuity. To diagnose how vCPU discontinuity causes the interference, the paper leverages NVIDIA Nsight Systems for fine-grained profiling and has the following findings: 1) vCPU discontinuity causes inefficient CPU-GPU synchronizations; 2) vCPU discontinuity delays task offloading to the vGPU; 3) Polling-based CPU-GPU synchronization suffers from interference more than blocking-based CPU-GPU synchronization; 4) GPU workloads with frequent task offloads and synchronizations are more vulnerable. Based on the findings, the paper proposes a novel polling-then-blocking CPU-GPU synchronization primitive. Evaluation shows that it can improve the performance by 4.2x.