Efficient GPU Computation Using Task Graph Parallelism

Efficient GPU Computation Using Task Graph Parallelism
复制标题

使用任务图并行性进行高效 GPU 计算

DOI:
10.1007/978-3-030-85665-6_27
复制
发表时间:
2021
期刊:
影响因子:
13.6
通讯作者:
Tsung
Tsung
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Dian;Tsung

文献摘要

被引文献

相似文献

最近,CUDA引入了一种新的任务图编程模型,CUDA图,以实现GPU工作的高效启动和执行。用户在任务图中描述GPU工作负载,而不是聚合GPU操作,允许CUDA运行时执行全图优化并显着降低内核调用开销。然而,编程CUDA图形是极具挑战性的。用户需要显式地构建带有详细参数设置的图,或者隐式地捕获需要使用流和事件进行复杂依赖关系和并发管理的图。为了克服这一挑战,我们引入了一个轻量级的任务图编程框架,以实现使用CUDA图的高效GPU计算。用户可以专注于依赖GPU操作的高级开发,而将流并发性和事件依赖性的所有复杂管理留给我们的优化算法。我们已经评估了我们的框架,并展示了其在微基准测试和大规模机器学习工作负载上的良好性能。结果还表明,我们的优化算法达到了与最佳构造图相当的性能,并且消耗的GPU资源少得多。
Recently, CUDA introduces a new task graph programming model,CUDA graph, to enable efficient launch and execution of GPU work. Users describe a GPU workload in a task graph rather than aggregated GPU operations, allowing the CUDA runtime to perform whole-graph optimization and significantly reduce the kernel call overheads. However, programming CUDA graphs is extremely challenging. Users need to explicitly construct a graph with verbose parameter settings or implicitly capture a graph that requires complex dependency and concurrency managements using streams and events. To overcome this challenge, we introduce a lightweight task graph programming framework to enable efficient GPU computation using CUDA graph. Users can focus on high-level development of dependent GPU operations, while leaving all the intricate managements of stream concurrency and event dependency to our optimization algorithm. We have evaluated our framework and demonstrated its promising performance on both micro-benchmarks and a large-scale machine learning workload. The result also shows that our optimization algorithm achieves very comparable performance to an optimally-constructed graph and consumes much less GPU resource.