LATTE-CC: Latency Tolerance Aware Adaptive Cache Compression Management for Energy Efficient GPUs

LATTE-CC: Latency Tolerance Aware Adaptive Cache Compression Management for Energy Efficient GPUs
复制标题

LATTE-CC:适用于节能 GPU 的延迟容忍感知自适应缓存压缩管理

DOI:
10.1109/hpca.2018.00028
复制
发表时间:
2018
期刊:
2018 IEEE International Symposium on High Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Carole
Carole
中科院分区:
--
文献类型:
--
作者:
A. Arunkumar;Shin;Vignesh Soundararajan;Carole

文献摘要

被引文献

相似文献

通用GPU应用程序受到内存子系统的效率和GPU上数据缓存容量可用性的严重限制。高速缓存压缩虽然能够扩展有效的高速缓存容量并提高高速缓存效率,但其代价是增加了命中延迟。这限制了高速缓存压缩应用于大多数较低级别的高速缓存,使其未被探索用于L1高速缓存和GPU。在GPU上直接应用最先进的高性能缓存压缩方案会导致从-52%到48%的广泛性能变化。为了最大限度地提高GPU缓存压缩的性能和能源效益,我们提出了一种新的压缩管理方案,称为LATTE-CC。LATTE-CC旨在利用GPU的动态变化延迟容限特性。LATTE-CC根据其对GPU流多处理器延迟容忍度的预测并通过在三种不同的压缩模式之间进行选择来压缩缓存行:无压缩、低延迟和高容量。LATTE-CC将高速缓存敏感的GPGPU应用程序的性能提高了48.4%,平均提高了19.2%,优于压缩算法的静态应用程序。LATTE-CC还将GPU能耗平均降低了10%,是最先进的压缩方案的两倍。
General-purpose GPU applications are significantly constrained by the efficiency of the memory subsystem and the availability of data cache capacity on GPUs. Cache compression, while is able to expand the effective cache capacity and improve cache efficiency, comes with the cost of increased hit latency. This has constrained the application of cache compression to mostly lower level caches, leaving it unexplored for L1 caches and for GPUs. Directly applying state-of-the-art high performance cache compression schemes on GPUs results in a wide performance variation from -52% to 48%. To maximize the performance and energy benefits of cache compression for GPUs, we propose a new compression management scheme, called LATTE-CC. LATTE-CC is designed to exploit the dynamically-varying latency tolerance feature of GPUs. LATTE-CC compresses cache lines based on its prediction of the degree of latency tolerance of GPU streaming multiprocessors and by choosing between three distinct compression modes: no compression, low-latency, and high-capacity. LATTE-CC improves the performance of cache sensitive GPGPU applications by as much as 48.4% and by an average of 19.2%, outperforming the static application of compression algorithms. LATTE-CC also reduces GPU energy consumption by an average of 10%, which is twice as much as that of the state-of-the-art compression scheme.