Principal Kernel Analysis: A Tractable Methodology to Simulate Scaled GPU Workloads

Principal Kernel Analysis: A Tractable Methodology to Simulate Scaled GPU Workloads
复制标题

DOI:
10.1145/3466752.3480100
复制
发表时间:
2021-10
期刊:
MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Cesar Avalos Baddouh;Mahmoud Khairy;Roland N. Green;Mathias Payer;Timothy G. Rogers
Cesar Avalos Baddouh;Mahmoud Khairy;Roland N. Green;Mathias Payer;Timothy G. Rogers
中科院分区:
其他
文献类型:
--
作者:
Cesar Avalos Baddouh;Mahmoud Khairy;Roland N. Green;Mathias Payer;Timothy G. Rogers

文献摘要

相似文献

在缩放的GPU工作负载中模拟所有线程会导致过高的模拟成本。周期级模拟比原生硅慢几个数量级,唯一的解决方案是减少模拟工作量,同时准确地表示程序。模拟GPU程序的现有解决方案要么缩放输入大小,模拟前几十亿条指令,要么模拟GPU和工作负载的一部分。这些解决方案缺乏对扩展系统的验证,产生不切实际的争用条件,并经常丢失关键代码部分。现有的CPU采样机制(如SimPoint)减少了每个线程的工作负载,并且不适合减少线程数量至关重要的GPU程序。GPU空间上的采样解决方案缺乏硅验证,需要按工作负载进行参数调整,并且无法扩展。需要一个易于处理的解决方案,在当代规模的工作负载上进行验证,以提供可靠的模拟结果。通过研究具有长达几个世纪的模拟时间的扩展工作负载,我们发现了现有解决方案的实际和算法局限性,并提出了主要内核分析:一种分层程序采样方法,通过使用可扩展的分析方法,易处理的聚类算法和内核内IPC稳定性检测来选择代表性的内核部分,从而简洁地表示GPU程序。我们使用X-Sim模拟器验证了147个工作负载和三代GPU的主要内核分析,证明了比以前工作更好的性能/错误权衡,并且长达一个世纪的MLPerf模拟减少到几个小时,平均周期误差为27%。
Simulating all threads in a scaled GPU workload results in prohibitive simulation cost. Cycle-level simulation is orders of magnitude slower than native silicon, the only solution is to reduce the amount of work simulated while accurately representing the program. Existing solutions to simulate GPU programs either scale the input size, simulate the first several billion instructions, or simulate a portion of both the GPU and the workload. These solutions lack validation against scaled systems, produce unrealistic contention conditions and frequently miss critical code sections. Existing CPU sampling mechanisms, like SimPoint, reduce per-thread workload, and are ill-suited to GPU programs where reducing the number of threads is critical. Sampling solutions on GPUs space lack silicon validation, require per-workload parameter tuning, and do not scale. A tractable solution, validated on contemporary scaled workloads, is needed to provide credible simulation results. By studying scaled workloads with centuries-long simulation times, we uncover practical and algorithmic limitations of existing solutions and propose Principal Kernel Analysis: a hierarchical program sampling methodology that concisely represents GPU programs by selecting representative kernel portions using a scalable profiling methodology, tractable clustering algorithm and detection of intra-kernel IPC stability. We validate Principal Kernel Analysis across 147 workloads and three GPU generations using the Accel-Sim simulator, demonstrating a better performance/error tradeoff than prior work and that century-long MLPerf simulations are reduced to hours with an average cycle error of 27% versus silicon.