Cooperative kernels: GPU multitasking for blocking algorithms

Cooperative kernels: GPU multitasking for blocking algorithms
复制标题

协作内核:用于阻塞算法的 GPU 多任务处理

DOI:
10.1145/3106237.3106265
复制
发表时间:
2017
期刊:
--
影响因子:
--
通讯作者:
Sorensen T
Sorensen T
中科院分区:
--
文献类型:
--
作者:
Sorensen T

文献摘要

参考文献

相似文献

人们对在 GPU 上加速不规则数据并行算法越来越感兴趣。这些算法通常是阻塞的,因此它们需要公平调度。但 GPU 编程模型(例如 OpenCL)并不要求公平调度,而且 GPU 调度程序在实践中也是不公平的。当前的方法通过利用当今 GPU 的调度怪癖,以不允许 GPU 与其他工作负载(例如图形渲染任务)共享的方式来避免此问题。我们提出了协作内核,这是传统 GPU 编程模型的扩展,旨在编写阻塞算法。协作内核的工作组是公平调度的,并且通过内核和调度程序协作的一小组语言扩展来支持多任务处理。我们描述了在 OpenCL 2.0 中实现的协作内核框架的原型实现,并通过将一组阻塞 GPU 应用程序移植到协作内核并检查它们在多任务处理下的性能来评估我们的方法。我们的原型没有利用特定于供应商的硬件、驱动程序或编译器支持,因此我们的结果提供了在实践中实现协作内核的效率下限。
There is growing interest in accelerating irregular data-parallel algorithms on GPUs. These algorithms are typically blocking, so they require fair scheduling. But GPU programming models (e.g. OpenCL) do not mandate fair scheduling, and GPU schedulers are unfair in practice. Current approaches avoid this issue by exploiting scheduling quirks of today's GPUs in a manner that does not allow the GPU to be shared with other workloads (such as graphics rendering tasks). We propose cooperative kernels, an extension to the traditional GPU programming model geared towards writing blocking algorithms. Workgroups of a cooperative kernel are fairly scheduled, and multitasking is supported via a small set of language extensions through which the kernel and scheduler cooperate. We describe a prototype implementation of a cooperative kernel framework implemented in OpenCL 2.0 and evaluate our approach by porting a set of blocking GPU applications to cooperative kernels and examining their performance under multitasking. Our prototype exploits no vendor-specific hardware, driver or compiler support, thus our results provide a lower-bound on the efficiency with which cooperative kernels can be implemented in practice.
DOI: 10.1109/surv.2012.021312.00045
发表时间: 2013-01-01
影响因子: 35.6
作者:
Vallina-Rodriguez, Narseo;Crowcroft, Jon
通讯作者: Crowcroft, Jon
多任务实时嵌入式 GPU 计算任务
DOI: --
发表时间: 2016
期刊: PMAM@PPoPP
影响因子: --
作者:
Pınar Muyan;John Douglas Owens
通讯作者: John Douglas Owens
C 中的协作多任务处理
DOI: --
发表时间: 1991
期刊:
影响因子: --
作者:
Marc Tarpenning
通讯作者: Marc Tarpenning
片段并行复合和过滤器
DOI: --
发表时间: 2010
期刊: Computer graphics forum (Print)
影响因子: --
作者:
Anjul Patney;Stanley Tzeng;John Douglas Owens
通讯作者: John Douglas Owens
在 GPU 加速器上迭​​代不规则 Maxflow 计算中利用并行性
DOI: --
发表时间: 2010
期刊: IEEE International Conference on High Performance Computing and Communications
影响因子: --
作者:
Steven Solomon;P. Thulasiraman;R. Thulasiram
通讯作者: R. Thulasiram