Near-Optimal Access Partitioning for Memory Hierarchies with Multiple Heterogeneous Bandwidth Sources

Near-Optimal Access Partitioning for Memory Hierarchies with Multiple Heterogeneous Bandwidth Sources
复制标题

具有多个异构带宽源的内存层次结构的近乎最优访问分区

DOI:
--
复制
发表时间:
2017
期刊:
International Symposium on High-Performance Computer Architecture
影响因子:
--
通讯作者:
S. Subramoney
S. Subramoney
中科院分区:
--
文献类型:
--
作者:
Jayesh Gaur;Mainak Chaudhuri;Pradeep Ramachandran;S. Subramoney

文献摘要

被引文献

相似文献

内存墙仍然是一个主要的性能瓶颈。虽然到目前为止,小型芯片上缓存在隐藏这一瓶颈方面是有效的,但现代应用程序的不断增加的内存占用使这种缓存变得无效。嵌入式DRAM(EDRAM)和高带宽存储器(HBM)等存储器技术的最新进展使大容量存储器能够集成到CPU封装上,作为除DDR主存储器之外的额外带宽来源。由于容量有限,这些存储器通常被实现为存储器端高速缓存。在传统智慧的推动下,许多旨在提高系统性能的优化都试图最大限度地提高内存端缓存的命中率。更高的命中率可以更好地利用缓存,因此被认为会带来更高的性能。在本文中,我们挑战了这一传统的观点,提出了一种动态访问分区算法DAP,它牺牲了缓存命中率来利用主存可用的未充分利用的带宽。DAP通过使用只需要16字节额外硬件的轻量级学习机制,在内存端缓存和主存之间实现了近乎最佳的带宽分区。仿真结果表明,当DAP在芯片堆叠的内存端DRAM缓存上实现时,平均性能提高了13%。我们还表明,DAP在不同的实现、带宽点和内存端高速缓存的容量点上提供了巨大的性能优势,使其成为任何基于片上SRAM高速缓存层次结构之外的多个异类带宽源的当前或未来系统的宝贵补充。
The memory wall continues to be a major performance bottleneck. While small on-die caches have been effective so far in hiding this bottleneck, the ever-increasing footprint of modern applications renders such caches ineffective. Recent advances in memory technologies like embedded DRAM (eDRAM) and High Bandwidth Memory (HBM) have enabled the integration of large memories on the CPU package as an additional source of bandwidth other than the DDR main memory. Because of limited capacity, these memories are typically implemented as a memory-side cache. Driven by traditional wisdom, many of the optimizations that target improving system performance have been tried to maximize the hit rate of the memory-side cache. A higher hit rate enables better utilization of the cache, and is therefore believed to result in higher performance. In this paper, we challenge this traditional wisdom and present DAP, a Dynamic Access Partitioning algorithm that sacrifices cache hit rates to exploit under-utilized bandwidth available at main memory. DAP achieves a near-optimal bandwidth partitioning between the memory-side cache and main memory by using a light-weight learning mechanism that needs just sixteen bytes of additional hardware. Simulation results show a 13% average performance gain when DAP is implemented on top of a die-stacked memory-side DRAM cache. We also show that DAP delivers large performance benefits across different implementations, bandwidth points, and capacity points of the memory-side cache, making it a valuable addition to any current or future systems based on multiple heterogeneous bandwidth sources beyond the on-chip SRAM cache hierarchy.