Analyzing and Leveraging Remote-Core Bandwidth for Enhanced Performance in GPUs

Analyzing and Leveraging Remote-Core Bandwidth for Enhanced Performance in GPUs
复制标题

DOI:
10.1109/pact.2019.00028
复制
发表时间:
2019-09
期刊:
2019 28th International Conference on Parallel Architectures and Compilation Techniques (PACT)
影响因子:
--
通讯作者:
M. Ibrahim;Hongyuan Liu;Onur Kayiran;Adwait Jog
M. Ibrahim;Hongyuan Liu;Onur Kayiran;Adwait Jog
中科院分区:
其他
文献类型:
--
作者:
M. Ibrahim;Hongyuan Liu;Onur Kayiran;Adwait Jog

文献摘要

被引文献

相似文献

从本地/共享缓存和内存中实现的带宽是图形处理单元(GPU)中的主要性能决定因素。这些现有的带宽来源通常不足以达到最佳的GPU性能。因此,为了进一步提高性能,我们专注于有效解锁带宽的其他潜在源,我们称之为遥控器带宽。该带宽的来源是基于这样的观察结果,即在其他GPU核心的局部(L1)CACHE中也可以找到一个GPU核心所需的数据(即L1读取误差)。在本文中,我们建议有效地协调跨GPU中核心的数据运动,以利用此遥远的带宽。但是,我们发现其有效的检测和利用提出了一些挑战。为此,我们专门讨论:a)跨核心共享哪些数据,b)核心具有共享数据,c)我们如何尽快获取数据。我们对广泛的GPGPU应用程序进行的广泛评估表明,由于从远程内核中获得的额外带宽,可以以适度的硬件成本来实现大幅改进的性能。
Bandwidth achieved from local/shared caches and memory is a major performance determinant in Graphics Processing Units (GPUs). These existing sources of bandwidth are often not enough for optimal GPU performance. Therefore, to enhance the performance further, we focus on efficiently unlocking an additional potential source of bandwidth, which we call as remote-core bandwidth. The source of this bandwidth is based on the observation that a fraction of data (i.e., L1 read misses) required by one GPU core can also be found in the local (L1) caches of other GPU cores. In this paper, we propose to efficiently coordinate the data movement across cores in GPUs to exploit this remote-core bandwidth. However, we find that its efficient detection and utilization presents several challenges. To this end, we specifically address: a) which data is shared across cores, b) which cores have the shared data, and c) how we can get the data as soon as possible. Our extensive evaluation across a wide set of GPGPU applications shows that significant performance improvement can be achieved at a modest hardware cost on account of the additional bandwidth received from the remote cores.