BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU Systems
BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU Systems
复制标题
DOI:
10.1109/isca52012.2021.00034
复制
发表时间:
2021-06
期刊:
影响因子:
--
通讯作者:
AmirAli Abdolrashidi;Hodjat Asghari Esfeden;A. Jahanshahi;Kaustubh Singh;N. Abu-Ghazaleh;Daniel Wong
中科院分区:
文献类型:
--
作者:
AmirAli Abdolrashidi;Hodjat Asghari Esfeden;A. Jahanshahi;Kaustubh Singh;N. Abu-Ghazaleh;Daniel Wong
As modern GPU workloads grow in size and complexity, there is an ever-increasing demand for GPU computational power. Emerging workloads contain hundreds or thousands of GPU kernel launches, which incur high overheads, and exhibit data-dependent behavior between kernels, which requires synchronization, leading to GPU under-utilization. Task-based execution models have been proposed to solve these issues, but they require significant programmer effort to port applications to proprietary task-based programming models in order to specify tasks and task dependencies. To address this need, we propose BlockMaestro, a software-hardware solution that combines command queue reordering, kernel-launch-time static analysis, and runtime hardware support to dynamically identify and resolve thread-block level data dependencies between kernels. Through static analysis of memory access patterns at kernel-launch-time, BlockMaestro can extract inter-kernel thread block-level data dependencies. BlockMaestro also introduces kernel pre-launching to reduce the kernel launch overheads experienced by multiple dependent kernels. Correctness is enforced by dynamically resolving thread block-level data dependency at runtime through hardware support. BlockMaestro achieves an average speedup of 51.76% (up to 2.92x) on data-dependent benchmarks, and requires minimal hardware overhead.