TicToc: Enabling Bandwidth-Efficient DRAM Caching for Both Hits and Misses in Hybrid Memory Systems

TicToc: Enabling Bandwidth-Efficient DRAM Caching for Both Hits and Misses in Hybrid Memory Systems
复制标题

TicToc:在混合内存系统中针对命中和未命中启用带宽高效的 DRAM 缓存

DOI:
10.1109/iccd46524.2019.00055
复制
发表时间:
2019
期刊:
2019 IEEE 37th International Conference on Computer Design (ICCD)
影响因子:
--
通讯作者:
Moinuddin K. Qureshi
Moinuddin K. Qureshi
中科院分区:
--
文献类型:
--
作者:
Vinson Young;Zeshan A. Chishti;Moinuddin K. Qureshi

文献摘要

被引文献

相似文献

本文研究了用于混合DRAM + 3D - XPoint存储器的高效带宽DRAM缓存。3D - XPoint作为一种能够实现大容量和非易失性主存系统的技术,正在成为DRAM的一种可行替代品。然而,3D - XPoint有几个特性限制了它完全取代DRAM:读取速度慢4 - 8倍,写入速度更慢。因此,在3D - XPoint之前有效的DRAM缓存对于实现大容量、低延迟和高写入带宽的存储器非常重要。目前,DRAM缓存设计主要有两种方法:(1)缓存行内标签(TIC)组织,通过在每行旁边存储标签,使得一次访问能同时获取标签和数据,从而针对命中进行优化;(2)缓存行外标签(TOC)组织,通过将来自多行数据的标签一起存储在一个标签行中,使得对一个标签行的一次访问能获取关于多行数据的信息,从而针对未命中进行优化。理想情况下,我们希望拥有TIC设计的低命中延迟以及TOC设计的低未命中带宽。为此,我们提出了一种TicToc组织,它同时具备TIC和TOC的功能,以获得两者在命中和未命中方面的优势。我们发现,简单地将这两种技术结合起来,实际上比单独使用TIC的性能更差,因为必须付出维护两种元数据的带宽成本。这项工作的主要贡献是开发了一些架构技术,以降低访问和维护TIC和TOC元数据的带宽成本。我们发现大部分更新带宽是由于维护TOC脏信息导致的。我们提出了一种DRAM缓存脏位技术,该技术将DRAM缓存脏信息传递到末级缓存,以帮助减少对已知脏行的重复脏位更新。我们还提出了一种抢先脏标记(PDM)技术,该技术预测哪些行将被写入,并在安装时主动标记脏位,以帮助避免对脏行的初始脏位更新。为了支持PDM,我们开发了一种新颖的基于程序计数器(PC)的写入预测器,以帮助仅标记可能被写入的行。我们对位于3D - XPoint之前的4GB DRAM缓存进行评估,结果表明我们的TicToc组织比基准TIC快10%,接近具有64MB SRAM标签的理想化DRAM缓存设计可能实现的14%的加速,而只需要34KB SRAM。
This paper investigates bandwidth-efficient DRAM caching for hybrid DRAM + 3D-XPoint memories. 3D-XPoint is becoming a viable alternative to DRAM as it enables high-capacity and non-volatile main memory systems. However, 3D-XPoint has several characteristics that limit it from outright replacing DRAM: 4-8x slower read, and even worse writes. As such, effective DRAM caching in front of 3D-XPoint is important to enable a high-capacity, low-latency, and high-write-bandwidth memory. There are currently two major approaches for DRAM cache design: (1) a Tag-Inside-Cacheline (TIC) organization that optimizes for hits, by storing tag next to each line such that one access gets both tag and data, and (2) a Tag-Outside-Cacheline (TOC) organization that optimizes for misses, by storing tags from multiple data lines together in a tag-line such that one access to a tag-line gets information on several data-lines. Ideally, we would like to have the low hit-latency of TIC designs, and the low miss-bandwidth of TOC designs. To this end, we propose a TicToc organization that provisions both TIC and TOC to get the hit and miss benefits of both. We find that naively combining both techniques actually performs worse than TIC individually, because one has to pay the bandwidth cost of maintaining both metadata. The main contribution of this work is developing architectural techniques to reduce bandwidth cost of accessing and maintaining both TIC and TOC metadata. We find that most of the update bandwidth is due to maintaining the TOC dirty information. We propose a DRAM Cache Dirtiness Bit technique that carries DRAM cache dirty information to last-level caches, to help prune repeated dirty-bit updates for known dirty lines. We also propose a Preemptive Dirty Marking (PDM) technique that predicts which lines will be written and proactively marks the dirty bit at install time, to help avoid the initial dirty-bit update for dirty lines. To support PDM, we develop a novel PC-based Write-Predictor to aid in marking only write-likely lines. Our evaluations on a 4GB DRAM cache in front of 3D-XPoint show that our TicToc organization enables 10% speedup over the baseline TIC, nearing the 14% speedup possible with an idealized DRAM cache design with 64MB of SRAM tags, while needing only 34KB SRAM.