BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU Systems

BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU Systems
复制标题

DOI:
10.1109/isca52012.2021.00034
复制
发表时间:
2021-06
期刊:
2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
AmirAli Abdolrashidi;Hodjat Asghari Esfeden;A. Jahanshahi;Kaustubh Singh;N. Abu-Ghazaleh;Daniel Wong
AmirAli Abdolrashidi;Hodjat Asghari Esfeden;A. Jahanshahi;Kaustubh Singh;N. Abu-Ghazaleh;Daniel Wong
中科院分区:
其他
文献类型:
--
作者:
AmirAli Abdolrashidi;Hodjat Asghari Esfeden;A. Jahanshahi;Kaustubh Singh;N. Abu-Ghazaleh;Daniel Wong

文献摘要

相似文献

随着现代 GPU 工作负载规模和复杂性的增长,对 GPU 计算能力的需求不断增加。新兴工作负载包含数百或数千个 GPU 内核启动,这会产生很高的开销,并且在内核之间表现出依赖于数据的行为,这需要同步,从而导致 GPU 利用率不足。基于任务的执行模型已经被提出来解决这些问题,但是它们需要程序员付出大量的努力才能将应用程序移植到专有的基于任务的编程模型,以便指定任务和任务依赖性。为了满足这一需求,我们提出了 BlockMaestro,这是一种软件硬件解决方案,它结合了命令队列重新排序、内核启动时静态分析和运行时硬件支持,以动态识别和解决内核之间的线程块级数据依赖关系。通过在内核启动时对内存访问模式进行静态分析,BlockMaestro 可以提取内核线程间块级数据依赖关系。 BlockMaestro 还引入了内核预启动,以减少多个相关内核所经历的内核启动开销。通过硬件支持在运行时动态解决线程块级数据依赖性来强制执行正确性。 BlockMaestro 在数据相关基准测试中实现了 51.76% 的平均加速(高达 2.92 倍),并且需要最少的硬件开销。
As modern GPU workloads grow in size and complexity, there is an ever-increasing demand for GPU computational power. Emerging workloads contain hundreds or thousands of GPU kernel launches, which incur high overheads, and exhibit data-dependent behavior between kernels, which requires synchronization, leading to GPU under-utilization. Task-based execution models have been proposed to solve these issues, but they require significant programmer effort to port applications to proprietary task-based programming models in order to specify tasks and task dependencies. To address this need, we propose BlockMaestro, a software-hardware solution that combines command queue reordering, kernel-launch-time static analysis, and runtime hardware support to dynamically identify and resolve thread-block level data dependencies between kernels. Through static analysis of memory access patterns at kernel-launch-time, BlockMaestro can extract inter-kernel thread block-level data dependencies. BlockMaestro also introduces kernel pre-launching to reduce the kernel launch overheads experienced by multiple dependent kernels. Correctness is enforced by dynamically resolving thread block-level data dependency at runtime through hardware support. BlockMaestro achieves an average speedup of 51.76% (up to 2.92x) on data-dependent benchmarks, and requires minimal hardware overhead.