Optimizing occupancy and ILP on the GPU using a combinatorial approach

Optimizing occupancy and ILP on the GPU using a combinatorial approach
复制标题

使用组合方法优化 GPU 上的占用率和 ILP

DOI:
10.1145/3368826.3377918
复制
发表时间:
2020
期刊:
Proceedings of the International Symposium on Code Generation and Optimization
影响因子:
--
通讯作者:
Mekhanoshin, Stanislav
Mekhanoshin, Stanislav
中科院分区:
--
文献类型:
--
作者:
Shobaki, Ghassan;Kerbow, Austin;Mekhanoshin, Stanislav

文献摘要

参考文献

被引文献

相似文献

本文提出了针对图形处理单元 (GPU) 编译时优化占用率和指令级并行性 (ILP) 问题的第一个通用解决方案。利用ILP(最小化调度长度)需要使用更多的寄存器,但使用更多的寄存器会降低占用率(可以并行运行的线程组的数量)。平衡这两个相互冲突的目标以实现最佳整体性能的问题是代码优化中一个具有挑战性的开放问题。在本文中,我们提出了一种双通道分支定界(B&B)算法,通过将占用率作为主要目标,将 ILP 作为次要目标来解决该问题。在第一遍中,算法搜索最大占用率时间表,而在第二遍中,算法迭代地搜索给出在第一遍中找到的最大占用率的最短时间表。所提出的调度算法在 LLVM 编译器中实现并应用于 AMD GPU。该算法的性能是使用 PlaidML 机器学习框架相对于 LLVM 的调度算法、AMD 的生产调度算法以及使用不同方法的现有 B&B 调度算法的基准进行评估的。结果表明,所提出的 B&B 调度算法相对于 LLVM 的调度器,几乎每个基准测试的速度提高了 35%,相对于 AMD 的调度器,提高了 31%,相对于现有的 B&B 调度器,提高了 18%。相对于 LLVM 的调度程序,几何平均改进为 16.3%,相对于 AMD 的生产调度程序,几何平均改进为 5.5%,相对于现有的 B&B 调度程序,几何平均改进为 6.2%。如果可以容忍更多的编译时间,相对于 AMD 的调度程序可以实现 6.3% 的几何平均改进。
This paper presents the first general solution to the problem of optimizing both occupancy and Instruction-Level Parallelism (ILP) when compiling for a Graphics Processing Unit (GPU). Exploiting ILP (minimizing schedule length) requires using more registers, but using more registers decreases occupancy (the number of thread groups that can be run in parallel). The problem of balancing these two conflicting objectives to achieve the best overall performance is a challenging open problem in code optimization. In this paper, we present a two-pass Branch-and-Bound (B&B) algorithm for solving this problem by treating occupancy as a primary objective and ILP as a secondary objective. In the first pass, the algorithm searches for a maximum-occupancy schedule, while in the second pass it iteratively searches for the shortest schedule that gives the maximum occupancy found in the first pass. The proposed scheduling algorithm was implemented in the LLVM compiler and applied to an AMD GPU. The algorithm’s performance was evaluated using benchmarks from the PlaidML machine learning framework relative to LLVM’s scheduling algorithm, AMD’s production scheduling algorithm and an existing B&B scheduling algorithm that uses a different approach. The results show that the proposed B&B scheduling algorithm speeds up almost every benchmark by up to 35% relative to LLVM’s scheduler, up to 31% relative to AMD’s scheduler and up to 18% relative to the existing B&B scheduler. The geometric-mean improvements are 16.3% relative to LLVM’s scheduler, 5.5% relative to AMD’s production scheduler and 6.2% relative to the existing B&B scheduler. If more compile time can be tolerated, a geometric-mean improvement of 6.3% relative to AMD’s scheduler can be achieved.
调度表达式 DAG 以满足最少的寄存器需求
DOI: --
发表时间: 1998
期刊: Computer languages
影响因子: --
作者:
Christoph W. Kessler
通讯作者: Christoph W. Kessler
使用组合优化方法实现寄存器压力最小化的预分配指令调度
DOI: 10.1145/2512432
发表时间: 2013
期刊: ACM Transactions on Architecture and Code Optimization (TACO)
影响因子: --
作者:
Ghassan Shobaki;Maxim Shawabkeh;Najm Eldeen Abu Rmaileh
通讯作者: Najm Eldeen Abu Rmaileh
最佳和启发式全局代码运动可最大限度地减少溢出
DOI: 10.1007/978-3-642-37051-9_2
发表时间: 2013
影响因子: 5.1
作者:
Gergö Barany;A. Krall
通讯作者: A. Krall
DOI: 10.1002/spe.2297
发表时间: 2015
期刊: Software: Practice and Experience
影响因子: --
作者:
Ghassan Shobaki;Laith Sakka;Najm Eldeen Abu Rmaileh;Hasan Al
通讯作者: Hasan Al
关联指令重新排序以减轻寄存器压力
DOI: --
发表时间: 2018
期刊: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
P. Rawat;Aravind Sukumaran;A. Rountev;F. Rastello;L. Pouchet;P. Sadayappan;P. Sadayappan
通讯作者: P. Sadayappan