SMGuard: A Flexible and Fine-Grained Resource Management Framework for GPUs

SMGuard: A Flexible and Fine-Grained Resource Management Framework for GPUs
复制标题

SMGuard:灵活且细粒度的 GPU 资源管理框架

DOI:
10.1109/tpds.2018.2848621
复制
发表时间:
2018-12
影响因子:
5.3
通讯作者:
Depei Qian
Depei Qian
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chao Yu;Yuebin Bai;Hailong Yang;Kun Cheng;Yuhao Gu;Zhongzhi Luan;Depei Qian

文献摘要

参考文献

被引文献

相似文献

GPU已经成为数据中心不可或缺的计算平台,将多个应用程序放在同一个GPU上广泛用于提高资源利用率。然而,由于不受控制的资源争用导致的性能干扰严重降低了协同定位应用的性能,并且无法提供满意的用户体验。在本文中,我们提出了一种软件方法SMGuard,它可以灵活地管理同一位置下多个应用程序的GPU资源使用情况。我们还提出了一种基于容量的GPU资源模型CapSM,该模型在协同定位的应用程序之间以细粒度的方式分配GPU资源。当延迟敏感型应用与批处理应用共存时,SMGuard可以通过配额机制防止批处理应用无限制地占用资源,通过预留机制保证延迟敏感型应用的资源使用。此外,SMGuard通过驱逐批处理应用的运行线程块来释放占用的资源,并将未完成的线程块重新映射到剩余资源,从而支持动态资源调整,避免了被抢占的内核重新启动。SMGuard是一种纯软件解决方案,不依赖于特殊的GPU硬件或编程模型,很容易在数据中心的商用GPU上采用。我们的评估显示,SMGuard在与批处理应用程序共存时,将延迟敏感型应用程序的平均性能提高了9.8倍。同时,GPU利用率平均可以提高35%。
GPUs have been becoming an indispensable computing platform in data centers, and co-locating multiple applications on the same GPU is widely used to improve resource utilization. However, performance interference due to uncontrolled resource contention severely degrades the performance of co-locating applications and fails to deliver satisfactory user experience. In this paper, we present SMGuard, a software approach to flexibly manage the GPU resource usage of multiple applications under co-location. We also propose a capacity based GPU resource model CapSM, which provisions the GPU resources in a fine-grained granularity among co-locating applications. When co-locating latency-sensitive applications with batch applications, SMGuard can prevent batch applications from occupying resources without constraint using quota based mechanism, and guarantee the resource usage of latency-sensitive applications with reservation based mechanism. In addition, SMGuard supports dynamic resource adjustment through evicting the running thread blocks of batch applications to release the occupied resources and remapping the uncompleted thread blocks to the remaining resources, which avoids the relaunch of the preempted kernel. The SMGuard is a pure software solution that does not rely on special GPU hardware or programming model, which is easy to adopt on commodity GPUs in data centers. Our evaluation shows that SMGuard improves the average performance of latency-sensitive applications by 9.8× when co-located with batch applications. In the meanwhile, the GPU utilization can be improved by 35 percent on average.
DOI: 10.1109/tpds.2016.2630697
发表时间: 2017-06
影响因子: 5.3
作者:
Yuhei Suzuki;Yusuke Fujii;Takuya Azumi;N. Nishio;S. Kato
通讯作者: Yuhei Suzuki;Yusuke Fujii;Takuya Azumi;N. Nishio;S. Kato
DOI: 10.1145/2872362.2872394
发表时间: 2016-03
期刊: Proceedings of the Twenty-First International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子: --
作者:
Haishan Zhu;M. Erez
通讯作者: Haishan Zhu;M. Erez
DOI: 10.1145/2925426.2926265
发表时间: 2016-06
期刊: Proceedings of the 2016 International Conference on Supercomputing
影响因子: --
作者:
Yunlong Xu;Rui Wang;Tao Li;Mingcong Song;Lan Gao;Zhongzhi Luan;D. Qian
通讯作者: Yunlong Xu;Rui Wang;Tao Li;Mingcong Song;Lan Gao;Zhongzhi Luan;D. Qian
DOI: 10.1109/aspdac.2014.6742976
发表时间: 2014-02
期刊: 2014 19th Asia and South Pacific Design Automation Conference (ASP-DAC)
影响因子: --
作者:
Paula Aguilera;Katherine Morrow;N. Kim
通讯作者: Paula Aguilera;Katherine Morrow;N. Kim
DOI: 10.1145/2000064.2000073
发表时间: 2011-06
期刊: 2011 38th Annual International Symposium on Computer Architecture (ISCA)
影响因子: --
作者:
Daniel Sánchez;Christos Kozyrakis
通讯作者: Daniel Sánchez;Christos Kozyrakis