Atos: A Task-Parallel GPU Scheduler for Graph Analytics

Atos: A Task-Parallel GPU Scheduler for Graph Analytics
复制标题

Atos:用于图形分析的任务并行 GPU 调度程序

DOI:
10.1145/3545008.3545056
复制
发表时间:
2022
期刊:
Proceedings of the 51st International Conference on Parallel Processing
影响因子:
--
通讯作者:
Owens, John
Owens, John
中科院分区:
--
文献类型:
--
作者:
Chen, Yuxin;Brock, Benjamin;Porumbescu, Serban;Buluc, Aydin;Yelick, Katherine;Owens, John

文献摘要

参考文献

被引文献

相似文献

提出了一种专门针对动态不规则应用的任务并行GPU动态调度框架ATOS。与主要的大容量同步并行(BSP)框架相比,ATOS通过支持依赖关系松散的应用程序的任务并行公式,实现了更高的GPU利用率,从而提供了额外的并发性,这对于并发瓶颈问题尤为重要。除了数据并行负载平衡之外,ATOS还提供隐式任务并行负载平衡,为用户提供了在它们之间进行平衡以实现最佳性能的灵活性。最后,ATOS允许用户通过控制内核策略和任务并行粒度来适应不同的用例。我们论证了这些控制在实践中的重要性。我们评估和分析了ATOS和BSP在三个应用上的性能:广度优先搜索、PageRank和图着色。与最先进的BSP GPU实施相比,ATOS实施在三个案例研究中实现了3.44倍、2.1倍和2.77倍的地理加速,以及12.8倍、3.2倍和9.08倍的峰值加速。除了简单地量化加速比之外,我们还广泛分析了每个加速比背后的原因。这种更深入的理解使我们能够得出如何为不同应用选择最佳ATOS配置的一般指导原则。最后,我们的分析为未来的动态调度框架设计提供了见解。
We present Atos, a task-parallel GPU dynamic scheduling framework that is especially targeted at dynamic irregular applications. Compared to the dominant Bulk Synchronous Parallel (BSP) frameworks, Atos exposes additional concurrency by supporting task-parallel formulations of applications with relaxed dependencies, achieving higher GPU utilization, which is particularly significant for problems with concurrency bottlenecks. Atos also offers implicit task-parallel load balancing in addition to data-parallel load balancing, providing users the flexibility to balance between them to achieve optimal performance. Finally, Atos allows users to adapt to different use cases by controlling the kernel strategy and task-parallel granularity. We demonstrate that each of these controls is important in practice.We evaluate and analyze the performance of Atos vs. BSP on three applications: breadth-first search, PageRank, and graph coloring. Atos implementations achieve geomean speedups of 3.44x, 2.1x, and 2.77x and peak speedups of 12.8x, 3.2x, and 9.08x across three case studies, compared to a state-of-the-art BSP GPU implementation. Beyond simply quantifying the speedup, we extensively analyze the reasons behind each speedup. This deeper understanding allows us to derive general guidelines for how to select the optimal Atos configuration for different applications. Finally, our analysis provides insights for future dynamic scheduling framework designs.
将面向集合的语言编译到大规模并行计算机上
DOI: 10.1109/fmpc.1988.47500
发表时间: 1988
期刊: Proceedings., 2nd Symposium on the Frontiers of Massively Parallel Computation
影响因子: --
作者:
G. Blelloch;G. Sabot
通讯作者: G. Sabot