BEAR: Techniques for mitigating bandwidth bloat in gigascale DRAM caches

BEAR: Techniques for mitigating bandwidth bloat in gigascale DRAM caches
复制标题

DOI:
10.1145/2749469.2750387
复制
发表时间:
2015-06
期刊:
2015 ACM/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Chiachen Chou;A. Jaleel;Moinuddin K. Qureshi
Chiachen Chou;A. Jaleel;Moinuddin K. Qureshi
中科院分区:
其他
文献类型:
--
作者:
Chiachen Chou;A. Jaleel;Moinuddin K. Qureshi

文献摘要

被引文献

相似文献

裸片堆叠存储器技术可实现千兆级DRAM高速缓存,其可在比商品DRAM高4倍至8倍的带宽下操作。当在该高速缓存中找到所请求的数据时,这种高速缓存可以通过以更快的速率服务数据来提高系统性能,从而潜在地将系统的存储器带宽增加4倍至8倍。不幸的是,DRAM高速缓存使用可用的存储器带宽不仅用于高速缓存命中时的数据传输,而且用于其他次要操作,例如高速缓存未命中检测、高速缓存未命中时的填充以及从最后一级片上高速缓存的脏逐出时的回写查找和内容更新。理想情况下,我们希望此类二次操作消耗的带宽可以忽略不计,并且几乎所有带宽都可用于将有用数据从DRAM缓存传输到处理器。我们评估了一个1GB的DRAM缓存,架构为合金缓存,并表明,即使是最带宽效率的建议DRAM缓存消耗3.8倍的带宽相比,一个理想化的DRAM缓存,不消耗任何带宽的二次操作。我们还表明,重新设计的DRAM缓存,以尽量减少二级操作所消耗的带宽可以潜在地提高系统性能的22%。为此,本文提出了带宽效率架构(BEAR)的DRAM缓存。BEAR集成了三个组件,每个组件用于减少未命中检测、未命中填充和回写探测所消耗的带宽。BEAR将DRAM缓存的带宽消耗降低了32%,从而将缓存命中延迟降低了24%,并将整体系统性能提高了10%。BEAR的开销可以忽略不计,优于理想化的SRAM标签存储设计,后者导致64兆字节的不可接受的开销,以及扇区缓存设计,后者导致6兆字节的SRAM存储开销。
Die stacking memory technology can enable gigascale DRAM caches that can operate at 4x-8x higher bandwidth than commodity DRAM. Such caches can improve system performance by servicing data at a faster rate when the requested data is found in the cache, potentially increasing the memory bandwidth of the system by 4x-8x. Unfortunately, a DRAM cache uses the available memory bandwidth not only for data transfer on cache hits, but also for other secondary operations such as cache miss detection, fill on cache miss, and writeback lookup and content update on dirty evictions from the last-level on-chip cache. Ideally, we want the bandwidth consumed for such secondary operations to be negligible, and have almost all the bandwidth be available for transfer of useful data from the DRAM cache to the processor. We evaluate a 1GB DRAM cache, architected as Alloy Cache, and show that even the most bandwidth-efficient proposal for DRAM cache consumes 3.8x bandwidth compared to an idealized DRAM cache that does not consume any bandwidth for secondary operations. We also show that redesigning the DRAM cache to minimize the bandwidth consumed by secondary operations can potentially improve system performance by 22%. To that end, this paper proposes Bandwidth Efficient ARchitecture (BEAR) for DRAM caches. BEAR integrates three components, one each for reducing the bandwidth consumed by miss detection, miss fill, and writeback probes. BEAR reduces the bandwidth consumption of DRAM cache by 32%, which reduces cache hit latency by 24% and increases overall system performance by 10%. BEAR, with negligible overhead, outperforms an idealized SRAM Tag-Store design that incurs an unacceptable overhead of 64 megabytes, as well as Sector Cache designs that incur an SRAM storage overhead of 6 megabytes.