Efficient GPU Computation Using Task Graph Parallelism
Efficient GPU Computation Using Task Graph Parallelism
复制标题
使用任务图并行性进行高效 GPU 计算
DOI:
10.1007/978-3-030-85665-6_27
复制
发表时间:
2021
期刊:
影响因子:
13.6
通讯作者:
Tsung
中科院分区:
文献类型:
--
作者:
Dian;Tsung
Recently, CUDA introduces a new task graph programming model,CUDA graph, to enable efficient launch and execution of GPU work. Users describe a GPU workload in a task graph rather than aggregated GPU operations, allowing the CUDA runtime to perform whole-graph optimization and significantly reduce the kernel call overheads. However, programming CUDA graphs is extremely challenging. Users need to explicitly construct a graph with verbose parameter settings or implicitly capture a graph that requires complex dependency and concurrency managements using streams and events. To overcome this challenge, we introduce a lightweight task graph programming framework to enable efficient GPU computation using CUDA graph. Users can focus on high-level development of dependent GPU operations, while leaving all the intricate managements of stream concurrency and event dependency to our optimization algorithm. We have evaluated our framework and demonstrated its promising performance on both micro-benchmarks and a large-scale machine learning workload. The result also shows that our optimization algorithm achieves very comparable performance to an optimally-constructed graph and consumes much less GPU resource.