Quality of service support for fine-grained sharing on GPUs

Quality of service support for fine-grained sharing on GPUs
复制标题

DOI:
10.1145/3079856.3080203
复制
发表时间:
2017-06
期刊:
2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Zhenning Wang;Jun Yang;R. Melhem;B. Childers;Youtao Zhang;M. Guo
Zhenning Wang;Jun Yang;R. Melhem;B. Childers;Youtao Zhang;M. Guo
中科院分区:
其他
文献类型:
--
作者:
Zhenning Wang;Jun Yang;R. Melhem;B. Childers;Youtao Zhang;M. Guo

文献摘要

被引文献

相似文献

GPU已被广泛应用于数据中心,为许多应用程序提供加速服务。共享GPU对于提高处理吞吐量和能效越来越重要。然而,并发应用程序之间的服务质量(QoS)是最低限度的支持。以前的努力是太粗粒度和不可扩展的QoS需求不断增加。我们提出了一个细粒度的GPU共享形式的QoS机制。我们的QoS支持可以在每个周期的基础上提供对内核进度的控制,以及每个内核的线程级并行量。由于精确的资源管理,我们的QoS支持与以前的最佳努力相比具有更好的可扩展性。评估表明,当GPU是由三个内核,其中两个有QoS目标共享,所提出的技术实现QoS目标43.8%,往往比以前的技术,并有20.5%的吞吐量高。
GPUs have been widely adopted in data centers to provide acceleration services to many applications. Sharing a GPU is increasingly important for better processing throughput and energy efficiency. However, quality of service (QoS) among concurrent applications is minimally supported. Previous efforts are too coarse-grained and not scalable with increasing QoS requirements. We propose QoS mechanisms for a fine-grained form of GPU sharing. Our QoS support can provide control over the progress of kernels on a per cycle basis and the amount of thread-level parallelism of each kernel. Due to accurate resource management, our QoS support has significantly better scalability compared with previous best efforts. Evaluations show that, when the GPU is shared by three kernels, two of which have QoS goals, the proposed techniques achieve QoS goals 43.8% more often than previous techniques and have 20.5% higher throughput.