Efficient and Fair Multi-programming in GPUs via Effective Bandwidth Management

Efficient and Fair Multi-programming in GPUs via Effective Bandwidth Management
复制标题

DOI:
10.1109/hpca.2018.00030
复制
发表时间:
2018-02
期刊:
2018 IEEE International Symposium on High Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Haonan Wang;Fan Luo;M. Ibrahim;Onur Kayiran;Adwait Jog
Haonan Wang;Fan Luo;M. Ibrahim;Onur Kayiran;Adwait Jog
中科院分区:
其他
文献类型:
--
作者:
Haonan Wang;Fan Luo;M. Ibrahim;Onur Kayiran;Adwait Jog

文献摘要

被引文献

相似文献

众所周知,通过限制GPGPU应用程序的线程级并行性(TLP)可以有效地改善整体性能。但是,我们发现,当将两个或多个应用程序共同在同一gpu上进行了两个或多个应用程序时,这种先前的技术可能会导致次优系统的吞吐量和公平性。这是因为他们试图孤立地最大化单个应用程序的性能,最终允许每个应用程序获得不成比例的共享资源。这导致共享缓存和内存中的高度争论。为了解决这个问题,我们为多应用执行环境提出了新的应用程序感知的TLP管理技术,以便所有共同的应用程序都可以善于明智地使用所有共享资源。为了测量这种使用,我们提出了一个称为有效带宽的应用级实用程序指标,该指标占了两个运行时指标:达到了DRAM带宽和Cache Miss率。我们发现,最大化总有效带宽并以平衡的方式在所有共同位置的应用中进行此操作可以显着改善系统吞吐量和公平性。与其详尽地搜索实现这些目标的TLP配置的所有不同组合,我们发现可以通过利用趋势来减少大量开销组合。我们提出的基于模式的TLP管理机制分别在基线上分别提高了20%和2倍的系统吞吐量和公平性,在该基线中,每个应用程序都以TLP配置执行,在单独执行时可以提供最佳性能。
Managing the thread-level parallelism (TLP) of GPGPU applications by limiting it to a certain degree is known to be effective in improving the overall performance. However, we find that such prior techniques can lead to sub-optimal system throughput and fairness when two or more applications are co-scheduled on the same GPU. It is because they attempt to maximize the performance of individual applications in isolation, ultimately allowing each application to take a disproportionate amount of shared resources. This leads to high contention in shared cache and memory. To address this problem, we propose new application-aware TLP management techniques for a multi-application execution environment such that all co-scheduled applications can make good and judicious use of all the shared resources. For measuring such use, we propose an application-level utility metric, called effective bandwidth, which accounts for two runtime metrics: attained DRAM bandwidth and cache miss rates. We find that maximizing the total effective bandwidth and doing so in a balanced fashion across all co-located applications can significantly improve the system throughput and fairness. Instead of exhaustively searching across all the different combinations of TLP configurations that achieve these goals, we find that a significant amount of overhead can be reduced by taking advantage of the trends, which we call patterns, in the way application's effective bandwidth changes with different TLP combinations. Our proposed pattern-based TLP management mechanisms improve the system throughput and fairness by 20% and 2x, respectively, over a baseline where each application executes with a TLP configuration that provides the best performance when it executes alone.