DyCache: Dynamic Multi-Grain Cache Management for Irregular Memory Accesses on GPU

DyCache: Dynamic Multi-Grain Cache Management for Irregular Memory Accesses on GPU
复制标题

DyCache:针对 GPU 上不规则内存访问的动态多粒度缓存管理

DOI:
10.1109/access.2018.2818193
复制
发表时间:
2018
期刊:
影响因子:
3.9
通讯作者:
Wang Zhiying
Wang Zhiying
中科院分区:
计算机科学3区
文献类型:
--
作者:
Guo Hui;Huang Libo;Lu Yashuai;Ma Sheng;Wang Zhiying

文献摘要

被引文献

相似文献

GPU利用宽高速缓存线(128 B)片上高速缓存为具有规则组织的数据结构的应用提供高带宽和高效的存储器访问。然而,新兴的应用程序表现出许多不规则的控制流和内存访问模式。不规则的内存访问会产生许多对L1数据缓存的细粒度内存访问。细粒度数据访问和粗粒度高速缓存设计之间的这种不匹配使得片上存储器空间更加受限,结果,高速缓存行替换的频率增加,并且L1数据高速缓存被低效地利用。提出细粒度缓存管理以提供高效的缓存管理,从而提高数据阵列的利用效率。与其他静态细粒度缓存管理不同,本文提出了一种动态多粒度缓存管理(DyCache),以解决一级数据缓存的低效使用问题。DyCache通过监控应用程序的访存模式,动态改变该高速缓存管理粒度,从而在不影响常规应用程序性能的前提下,提高非常规访存应用程序的GPU性能。我们的实验表明,DyCache可以实现40%的几何平均改善IPC与基线缓存(128 B)的不规则内存访问的应用程序,而对于有规律的内存访问的应用程序,DyCache不会降低性能。
GPU utilizes the wide cache-line (128B) on-chip cache to provide high bandwidth and efficient memory accesses for applications with regularly-organized data structures. However, emerging applications exhibit a lot of irregular control flows and memory access patterns. Irregular memory accesses generate many fine-grain memory accesses to L1 data cache. This mismatching between fine-grain data accesses and the coarse-grain cache design makes the on-chip memory space more constrained and as a result, the frequency of cache line replacement increases and L1 data cache is utilized inefficiently. Fine-grain cache management is proposed to provide efficient cache management to improve the efficiency of data array utilization. Unlike other static fine-grain cache managements, we propose a dynamic multi-grain cache management, called DyCache, to resolve the inefficient use of L1 data cache. Through monitoring the memory access pattern of applications, DyCache can dynamically alter the cache management granularity in order to improve the performance of GPU for applications with irregular memory accesses while not impact the performance for regular applications. Our experiment demonstrates that DyCache can achieve a 40% geometric mean improvement on IPC for applications with irregular memory accesses against the baseline cache (128B), while for applications with regular memory accesses, DyCache does not degrade the performance.