Optimizing occupancy and ILP on the GPU using a combinatorial approach
Optimizing occupancy and ILP on the GPU using a combinatorial approach
复制标题
使用组合方法优化 GPU 上的占用率和 ILP
DOI:
10.1145/3368826.3377918
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Mekhanoshin, Stanislav
中科院分区:
文献类型:
--
作者:
Shobaki, Ghassan;Kerbow, Austin;Mekhanoshin, Stanislav
This paper presents the first general solution to the problem of optimizing both occupancy and Instruction-Level Parallelism (ILP) when compiling for a Graphics Processing Unit (GPU). Exploiting ILP (minimizing schedule length) requires using more registers, but using more registers decreases occupancy (the number of thread groups that can be run in parallel). The problem of balancing these two conflicting objectives to achieve the best overall performance is a challenging open problem in code optimization. In this paper, we present a two-pass Branch-and-Bound (B&B) algorithm for solving this problem by treating occupancy as a primary objective and ILP as a secondary objective. In the first pass, the algorithm searches for a maximum-occupancy schedule, while in the second pass it iteratively searches for the shortest schedule that gives the maximum occupancy found in the first pass. The proposed scheduling algorithm was implemented in the LLVM compiler and applied to an AMD GPU. The algorithm’s performance was evaluated using benchmarks from the PlaidML machine learning framework relative to LLVM’s scheduling algorithm, AMD’s production scheduling algorithm and an existing B&B scheduling algorithm that uses a different approach. The results show that the proposed B&B scheduling algorithm speeds up almost every benchmark by up to 35% relative to LLVM’s scheduler, up to 31% relative to AMD’s scheduler and up to 18% relative to the existing B&B scheduler. The geometric-mean improvements are 16.3% relative to LLVM’s scheduler, 5.5% relative to AMD’s production scheduler and 6.2% relative to the existing B&B scheduler. If more compile time can be tolerated, a geometric-mean improvement of 6.3% relative to AMD’s scheduler can be achieved.
登录
查看更多内容
DOI:
--
发表时间:
1998
期刊:
Computer languages
影响因子:
--
作者:
Christoph W. Kessler
通讯作者:
Christoph W. Kessler
DOI:
10.1145/2512432
发表时间:
2013
期刊:
ACM Transactions on Architecture and Code Optimization (TACO)
影响因子:
--
作者:
Ghassan Shobaki;Maxim Shawabkeh;Najm Eldeen Abu Rmaileh
通讯作者:
Najm Eldeen Abu Rmaileh
影响因子:
5.1
作者:
Gergö Barany;A. Krall
通讯作者:
A. Krall
DOI:
10.1002/spe.2297
发表时间:
2015
期刊:
Software: Practice and Experience
影响因子:
--
作者:
Ghassan Shobaki;Laith Sakka;Najm Eldeen Abu Rmaileh;Hasan Al
通讯作者:
Hasan Al
DOI:
--
发表时间:
2018
期刊:
International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
作者:
P. Rawat;Aravind Sukumaran;A. Rountev;F. Rastello;L. Pouchet;P. Sadayappan;P. Sadayappan
通讯作者:
P. Sadayappan