DrGPU: A Top-Down Profiler for GPU Applications

DrGPU: A Top-Down Profiler for GPU Applications
复制标题

DOI:
10.1145/3578244.3583736
复制
发表时间:
2023-04
期刊:
Proceedings of the 2023 ACM/SPEC International Conference on Performance Engineering
影响因子:
--
通讯作者:
Yueming Hao;Nikhil Jain;R. Van der Wijngaart;N. Saxena;Yuanbo Fan;Xu Liu
Yueming Hao;Nikhil Jain;R. Van der Wijngaart;N. Saxena;Yuanbo Fan;Xu Liu
中科院分区:
其他
文献类型:
--
作者:
Yueming Hao;Nikhil Jain;R. Van der Wijngaart;N. Saxena;Yuanbo Fan;Xu Liu

文献摘要

被引文献

相似文献

GPU 在 HPC 系统中已变得很常见,可加速科学计算和机器学习应用程序。有效地将这些应用程序映射到快速发展的 GPU 架构以实现高性能是一项众所周知的挑战。 GPU 内核中存在各种性能低下问题,阻碍应用程序获得裸机性能。虽然现有工具能够衡量这些低效率,但它们主要侧重于数据收集和呈现,需要大量的手动工作来了解可操作优化的根本原因。因此,我们开发了 DrGPU,这是一种新颖的分析器,可以执行自上而下的分析来指导 GPU 代码优化。作为其显着特征,DrGPU 利用商用 GPU 中提供的硬件性能计数器来量化停顿周期,将其分解为各种停顿原因,查明根本原因,并提供直观的优化指导。在 DrGPU 的帮助下,我们能够分析重要的 GPU 基准测试和应用程序,并获得显着的加速 — V100 上高达 1.77 倍,GTX 1650 上高达 2.03 倍。
GPUs have become common in HPC systems to accelerate scientific computing and machine learning applications. Efficiently mapping these applications to rapid evolutions of GPU architectures for high performance is a well-known challenge. Various performance inefficiencies exist in GPU kernels that impede applications from obtaining bare-metal performance. While existing tools are able to measure these inefficiencies, they mostly focus on data collection and presentation, requiring significant manual efforts to understand the root causes for actionable optimization. Thus, we develop DrGPU, a novel profiler that performs top-down analysis to guide GPU code optimization. As its salient feature, DrGPU leverages hardware performance counters available in commodity GPUs to quantify stall cycles, decompose them into various stall reasons, pinpoint root causes, and provide intuitive optimization guidance. With the help of DrGPU, we are able to analyze important GPU benchmarks and applications and obtain nontrivial speedups --- up to 1.77X on V100 and 2.03X on GTX 1650.