Quality of service support for fine-grained sharing on GPUs
Quality of service support for fine-grained sharing on GPUs
复制标题
DOI:
10.1145/3079856.3080203
复制
发表时间:
2017-06
期刊:
影响因子:
--
通讯作者:
Zhenning Wang;Jun Yang;R. Melhem;B. Childers;Youtao Zhang;M. Guo
中科院分区:
文献类型:
--
作者:
Zhenning Wang;Jun Yang;R. Melhem;B. Childers;Youtao Zhang;M. Guo
GPUs have been widely adopted in data centers to provide acceleration services to many applications. Sharing a GPU is increasingly important for better processing throughput and energy efficiency. However, quality of service (QoS) among concurrent applications is minimally supported. Previous efforts are too coarse-grained and not scalable with increasing QoS requirements. We propose QoS mechanisms for a fine-grained form of GPU sharing. Our QoS support can provide control over the progress of kernels on a per cycle basis and the amount of thread-level parallelism of each kernel. Due to accurate resource management, our QoS support has significantly better scalability compared with previous best efforts. Evaluations show that, when the GPU is shared by three kernels, two of which have QoS goals, the proposed techniques achieve QoS goals 43.8% more often than previous techniques and have 20.5% higher throughput.