Hierarchical Roofline analysis for GPUs: Accelerating performance optimization for the NERSC‐9 Perlmutter system

Hierarchical Roofline analysis for GPUs: Accelerating performance optimization for the NERSC‐9 Perlmutter system
复制标题

GPU 的分层 Roofline 分析:加速 NERSC-9 Perlmutter 系统的性能优化

DOI:
--
复制
发表时间:
2019
期刊:
Concurrency and Computation
影响因子:
--
通讯作者:
Samuel Williams
Samuel Williams
中科院分区:
--
文献类型:
--
作者:
Charlene Yang;T. Kurth;Samuel Williams

文献摘要

被引文献

相似文献

车顶性能模型提供了一种直观且有洞察力的方法来识别性能瓶颈并指导性能优化。为了为NERSC的下一代超级计算机Perlmuter做准备,本文提出了一种在NVIDIA图形处理器上构建分层屋顶线的方法,并对其进行了扩展,以支持降低的精度和张量内核。分层Roofline将L1、L2、设备内存和系统内存带宽整合到一个图形中,与仅使用DRAM的传统Roofline相比,它为性能分析提供了更深刻的见解。我们使用Roofline方法分析了三个代理应用程序:来自BerkeleyGW的GPP、来自Amrex的HPGMG和来自TensorFlow的Conv2d。通过这样做,我们展示了我们的方法能够容易地了解NVIDIA图形处理器上的性能和性能瓶颈的各个方面,并促进代码优化。
The Roofline performance model provides an intuitive and insightful approach to identifying performance bottlenecks and guiding performance optimization. In preparation for the next‐generation supercomputer Perlmutter at NERSC, this paper presents a methodology to construct a hierarchical Roofline on NVIDIA GPUs and extends it to support reduced precision and Tensor Cores. The hierarchical Roofline incorporates L1, L2, device memory, and system memory bandwidths into one single figure, and it offers more profound insights into performance analysis than the traditional DRAM‐only Roofline. We use our Roofline methodology to analyze three proxy applications: GPP from BerkeleyGW, HPGMG from AMReX, and conv2d from TensorFlow. In doing so, we demonstrate the ability of our methodology to readily understand various aspects of performance and performance bottlenecks on NVIDIA GPUs and motivate code optimizations.