QoS-aware dynamic resource allocation for spatial-multitasking GPUs

QoS-aware dynamic resource allocation for spatial-multitasking GPUs
复制标题

DOI:
10.1109/aspdac.2014.6742976
复制
发表时间:
2014-02
期刊:
2014 19th Asia and South Pacific Design Automation Conference (ASP-DAC)
影响因子:
--
通讯作者:
Paula Aguilera;Katherine Morrow;N. Kim
Paula Aguilera;Katherine Morrow;N. Kim
中科院分区:
其他
文献类型:
--
作者:
Paula Aguilera;Katherine Morrow;N. Kim

文献摘要

被引文献

相似文献

基于gpu的通用计算(GPGPU computing)正被广泛采用;但是,有些GPGPU应用不能充分利用GPU资源。在这些情况下,空间多任务通过在同时运行的应用程序之间划分GPU资源,更好地利用GPU提供的并行性。当一个或多个这样的应用程序有服务质量(QoS)需求时,必须为这些应用程序分配足够的资源来满足它们的需求。可以禁用剩余资源以减少功耗,也可以将其用于加速其他应用程序。然而,我们观察到QoS应用程序满足其性能需求的资源量部分取决于协同执行的应用程序。在本文中,我们提出了一种运行时技术,在并发运行的应用程序之间动态划分GPU资源-至少其中一个具有QoS要求。我们证明了所提出的技术可以满足100%的QoS要求,同时还实现了7W的功耗降低或17.57%的性能改进,以共同执行尽力而为的应用程序。
General-purpose computing on GPUs (GPGPU computing) is becoming widely adopted; however, some GPGPU applications fail to fully utilize GPU resources. In these cases, spatial multitasking better exploits the parallelism offered by GPUs by partitioning the GPU resources among simultaneously-running applications. When one or more such applications have quality-of-service (QoS) requirements, enough resources must be allocated for those applications to satisfy their requirements. Remaining resources can be either disabled to reduce power consumption or used to accelerate other applications. However, we observe that the amount of resources for a QoS application to satisfy its performance requirement is dependent in part upon the co-executing applications. In this paper, we propose a runtime technique to dynamically partition GPU resources between concurrently running applications - at least one of which has a QoS requirement. We demonstrate that the proposed technique can satisfy a 100% QoS requirement while also achieving either a 7W power consumption reduction or a 17.57% performance improvement for co-executing best-effort applications.