Adaptively Reduced DRAM Caching for Energy-Efficient High Bandwidth Memory

Adaptively Reduced DRAM Caching for Energy-Efficient High Bandwidth Memory
复制标题

DOI:
10.1109/tc.2022.3140897
复制
发表时间:
2022-10
影响因子:
3.7
通讯作者:
Payman Behnam;M. N. Bojnordi
Payman Behnam;M. N. Bojnordi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Payman Behnam;M. N. Bojnordi

文献摘要

被引文献

相似文献

封装内DRAM缓存提供了比传统存储系统更高的带宽。使缓存管理适应每个应用程序的运行时特征似乎是提高带宽效率和性能的一种有前途的方法。遗憾的是,细粒度的缓存块监控和适配往往由于其显著的带宽、性能和硬件开销而变得不切实际。提出了一种使用两个在运行时可调的参数来监控缓存块的新机制。我们提出了两种低成本的基于计数器的机制来实现DRAM中的块监控。此外,我们提出了一种新的调度机制,当数据移动开销达到最小时,该机制将计数器信息机会性地传输到DRAM堆栈。我们在一组数据密集型并行应用上的模拟结果表明,与现有的DRAM缓存结构相比,该机制的性能平均提高了31%和24%。相同基准下的系统节能效果平均为29%-18%。
In-package DRAM cache provides a higher bandwidth than conventional memory systems. Adapting the cache management to the run-time characteristics of each application seems a promising approach improving bandwidth efficiency and performance. Regrettably, fine-grained cache block monitoring and adaptation often becomes impractical due to its significant bandwidth, performance and hardware overheads. This paper proposes a novel mechanism for monitoring cache blocks using two parameters that are adjustable at run time. We propose two low-cost counter-based mechanisms to realize the block monitors in DRAM. Moreover, we propose a novel scheduling mechanism that opportunistically transfers the counter information to the DRAM stack when the data movement overhead reaches its minimum. Our simulation results on a set of data intensive parallel applications indicate that the proposed mechanisms achieve averages of 31%, 24% performance improvements over the state-of-the-art DRAM cache architectures. System energy savings over the same baselines are 29%, 18% on average.